{"api_version":"v1","generated_at":"2026-09-29T13:00:45","count":772,"scope":"latest","window_days":14,"since":"2026-09-10","papers":[{"id":"2609.30563","version":1,"title":"Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content","zh_title":"少思考以更好地模拟：直觉提示提升LLM智能体模拟个体社交媒体反应（包括不熟悉内容）","abstract":"Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.","authors":["Ljubisa Bojic","Tijana Stanic","Joerg Matthes","Agariadne Dwinggo Samala","Bojana Dinic","Jue Wang"],"categories":["cs.AI","cs.CL","cs.HC","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30563","pdf_url":"https://arxiv.org/pdf/2609.30563","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","提示策略"],"reason":"用LLM预测真实个体社交媒体反应，与人类数据对照，评估提示策略对仿真保真度的影…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":1,"question":"在社交媒体反应预测中，不同的提示设计（档案内容与指令风格）如何影响大语言模型模拟特定个体反应的保真度？","design":"用四个大语言模型扮演八名塞尔维亚参与者，基于问卷、深度访谈和书面自我呈现构建个人档案，在五种提示条件下（变化档案内容和指令风格，如人口学背景、态度内容、直觉式回应等）预测他们对68条社交媒体帖子的反应，测量预测与真实反应的一致性。","baseline":"八名塞尔维亚参与者的真实社交媒体反应数据，以及他们自己的问卷回答（用于比较自我一致性）。","findings":"态度性档案内容比人口学背景大幅提升预测准确度；直觉式指令（要求模型凭直觉立即回应而非分析式推理）在所有条件中保真度最高，并将个体差异压缩从人类的七倍降至三倍。","reliability":"论文未讨论","relevance":"该研究直接检验LLM模拟个体行为的保真度，与人类真实反应对照，并揭示提示策略对仿真偏差的影响，对评估LLM作为人类被试替代品的可靠性具有关键参考价值。","inspiration":"值得借鉴的是通过改变提示指令风格（直觉式 vs 分析式）来操纵模型的认知模式，并测量其对个体差异保真度的影响｜可迁移到消费者金融决策实验，如模拟个体在信贷选择或储蓄行为中的异质性反应｜用LLM扮演不同风险偏好和金融素养的消费者，处理为直觉式或分析式提示，结果变量为信贷产品选择或投资决策，与真实消费者调查或实验数据对照。"}},{"id":"2609.30883","version":1,"title":"Warned alike, AI agents avoid the less-crowded road while people take it","zh_title":"同样被警告，AI智能体避开较不拥挤的道路而人类选择它","abstract":"AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switching alone. The warning discouraged the very move it predicted. The pattern persisted for 100 rounds. Two other model families shifted the same way without locking onto one road. Twelve all-human groups (240 participants) stayed near balance under numerical reports or the tip and warning. In 24 mixed groups with a further 240 participants, imbalance grew with the share of agents in the registered analysis, while people increasingly took the road the agents avoided. Collective costs stayed below the allagent reference, but with 15 agents and 5 humans, agent seats averaged 80 min, compared with 44 min for human seats. Shared forecasts can thus sustain collective inefficiency among similar agents. A better group average can also hide an unequal burden. Evaluations of AI agents that share resources should test populations, treat messages as interventions and report who bears the costs.","authors":["Takahiro Ezaki","Naoto Imura","Katsuhiro Nishinari"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30883","pdf_url":"https://arxiv.org/pdf/2609.30883","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","拥堵博弈","人机对照"],"reason":"用GPT agent群体模拟拥堵博弈，并与240名人类被试对照，发现警告导致a…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":2,"question":"共享预测信息（警告）是否会导致AI智能体在拥堵博弈中持续选择拥挤道路，从而造成集体低效？","design":"用50个GPT智能体（gpt-5.4-mini）模拟通勤者，在双路径拥堵博弈中，操纵每日广播信息（无报告、数值报告、提示、提示+警告），测量路径不平衡度、切换比例和平均旅行时间。","baseline":"12个全人类组（240名参与者）在相同博弈和广播条件下的行为数据，以及24个混合组（240名参与者）与智能体共存时的行为。","findings":"警告导致GPT智能体群体持续拥挤一条道路，平均旅行时间从64分钟升至95分钟，尽管个体切换可节省69分钟；全人类组保持接近均衡，混合组中人类更多选择智能体避开的道路，且智能体承担更高成本。","reliability":"论文指出共享预测可导致相似智能体间的集体低效，且群体平均成本可能掩盖负担不均；评估共享资源的AI智能体应测试群体、将消息视为干预并报告成本承担者。","relevance":"高度相关：该研究用LLM智能体模拟人类在拥堵博弈中的决策，并与真实人类数据对照，揭示了共享信息对集体行为的负面影响，直接回应了仿真可靠性与偏差问题。","inspiration":"值得借鉴的是将公共信息（警告）作为干预变量，观察其对群体决策动态的影响，并设置全人类和混合组对照以分离智能体特有行为。｜可迁移到政策公告的预期形成场景，如央行沟通对金融市场参与者行为的影响。｜设计：用LLM智能体模拟投资者，处理为央行发布的不同措辞的前瞻指引，结果变量为资产配置集中度和市场波动率，对照真实投资者在类似公告下的交易数据。"}},{"id":"2609.30896","version":1,"title":"Large language models underestimate and partly misrepresent cultural variation in everyday norms","zh_title":"大语言模型低估并部分误现日常规范的文化差异","abstract":"A key aspect of culture is a society's norms about everyday behavior. How accurately do large language models (LLMs) represent cultural differences in such norms? To answer this question we used the recent Global Study of Everyday Norms (GSEN), which collected ratings of 150 scenarios in 90 societies, as the human benchmark. We prompted GPT-5 to estimate each society's average rating for every scenario, and later repeated the benchmark in three other LLMs: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Compared to GSEN estimates, all four LLMs misrepresented cultural variation in two ways. First, they greatly underestimated its magnitude, estimating differences between societies to be, on average, less than half their measured size. Second, for many scenarios the LLMs poorly identified the pattern of variation, that is, which societies judged the behavior less acceptable and which societies judged it more acceptable. The pattern of variation was identified better for scenarios that elicit concerns about vulgarity, especially scenarios involving kissing and flirting. We also found that norms in more developed societies tended to be estimated somewhat more accurately, and that prompting in local survey languages rather than English produced only a modest improvement in accuracy. Local-language prompting also reduced, but did not remove, the underestimation of between-society differences. Cultural differences in everyday norms are only weakly and unevenly represented by LLMs.","authors":["Kimmo Eriksson","Irina Vartanova","Pontus Strimling"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30896","pdf_url":"https://arxiv.org/pdf/2609.30896","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化规范","仿真偏差","人类数据对照"],"reason":"用LLM估计社会规范并与真实人类调查数据对照，评估仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":3,"question":"大语言模型能否准确表示不同社会之间日常行为规范的文化差异？","design":"用 GPT-5、GPT-5.4、Claude Opus 4.6 和 Gemini 3.1 Pro 四个 LLM，以英语或当地语言提示，估计 90 个社会对 150 个日常行为场景的平均可接受度评分，并与 GSEN 调查数据比较。","baseline":"全球日常规范研究（GSEN）中 90 个社会超过 25,000 名参与者对 150 个场景的评分。","findings":"四个 LLM 都大幅低估了社会间规范差异的幅度，估计的差异平均不到实际测量的一半；对于许多场景，LLM 识别差异模式的能力较差，但在涉及粗俗（如亲吻和调情）的场景上表现较好。","reliability":"论文指出 LLM 对较发达社会的规范估计更准确，用当地语言提示仅带来适度改进，且未能消除对社会间差异的低估；LLM 对文化差异的表示既弱且不均衡。","relevance":"该研究直接评估 LLM 在跨文化社会规范仿真中的偏差，与研究者关注的人类仿真可靠性及失效条件高度契合，值得精读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其用真实大规模跨国调查作为基准、将仿真误差分解为幅度和模式两个维度的方法，并检验提示语言等处理变量的影响。｜可迁移到跨国消费者行为或政策接受度的仿真研究，例如不同国家居民对环保政策、金融产品条款或广告伦理的接受度差异。｜以 LLM 模拟不同国家被试，对一系列政策或产品场景给出接受度评分，处理变量为提示语言（英语 vs 当地语言），结果变量为评分，与真实跨国调查数据（如世界价值观调查或特定政策民意调查）对照，评估 LLM 仿真的幅度压缩和模式偏差。"}},{"id":"2604.20050","version":4,"title":"Information Aggregation with AI Agents","zh_title":"AI代理的信息聚合研究","abstract":"Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy information structures, performance deteriorates in the harder structures, suggesting that AI agents struggle in environments where more than two levels of interactive reasoning are required, a ceiling close to the one documented in human subjects. Consistent with our theoretical predictions, market accuracy does not improve from allowing cheap talk communication, changing the duration of the market, or strategic prompting; initial price has little average effect but matters in the very hard structure. We also find that ``smarter'' AI agents perform better at aggregation and are more profitable. Surprisingly, giving them feedback about past performance does not improve aggregation. A further wave of markets, run three months later with capability-frontier models, aggregates information more often in the three easier structures but not in the hardest one, where higher capability replaces markets that are confidently wrong with markets that hedge near 0.5.","authors":["Spyros Galanis"],"categories":["econ.GN","cs.AI","cs.GT","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-04-21","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2604.20050","pdf_url":"https://arxiv.org/pdf/2604.20050","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","信息聚合","行为实验"],"reason":"用AI代理模拟人类交易行为，并与人类被试结果对照，评估信息聚合能力。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":4,"question":"AI 代理能否通过交易聚合分散的私人信息，并像人类一样通过观察价格变动推断他人知识？","design":"用八种大语言模型（Claude Haiku 3.5/4.5、Gemini 2.5/3 Flash、GPT-4o/5 mini、gemma3:4b、qwen3:8b）组成三人交易团队，在四种难度递增的信息结构中交易预测市场证券；处理包括允许廉价谈话、战略提示、初始价格（0.3/0.5/0.7）、市场时长（3/6/9轮）和反馈信息，共144种配置，每种至少运行12次，生成1772个市场；结果变量为市场准确率（价格是否接近真实价值）和利润。","baseline":"人类被试在类似信息结构中的推理层级上限（约两级交互推理），来自已有实验文献。","findings":"在简单信息结构中，中位数市场能有效聚合信息，但在需要超过两级交互推理的困难结构中表现恶化；允许廉价谈话、改变市场时长或战略提示均未提高市场准确率，初始价格平均影响小但在极难结构中重要。更聪明的AI代理聚合更好且更盈利，但反馈过去表现未改善聚合；三个月后用前沿模型重跑，在三个较易结构中聚合更频繁，但在最难结构中未改善，高能力模型将自信错误的市场替换为在0.5附近对冲的市场。","reliability":"论文承认AI代理在需要超过两级交互推理的环境中挣扎，且能力提升并未解决最难结构中的聚合失败；未明确讨论其他失效条件。","relevance":"该研究直接以AI代理模拟人类交易者，并与人类被试的推理层级上限对照，评估信息聚合能力，属于用LLM进行人类仿真实验并检验可靠性的核心工作，值得精读。","inspiration":"借鉴其系统操纵信息结构复杂度、市场机制参数（时长、初始价格、沟通）和模型能力来测量聚合效率与利润的做法，并设置理论基准（可分离证券的完全聚合预测）｜可迁移到资产定价实验中的信息效率研究，如内幕交易监管、分析师预测市场或央行沟通对价格发现的影响｜用不同能力LLM作为交易者，在预测市场中交易与真实宏观经济指标挂钩的证券，处理为是否允许公开评论或改变交易轮次，结果变量为价格偏离真实值的程度，并与人类实验数据（如Plott & Sunder的经典信息聚合实验）对照。"}},{"id":"2609.02729","version":2,"title":"BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research","zh_title":"BuildOcc：用于建筑能源研究的大语言模型居住者智能体平台","abstract":"Occupants are a primary source of uncertainty in building energy consumption and management, yet existing occupant behavior models cannot capture adaptive and reasoning responses considering the occupant's personal history, current context, and the type of energy signal being delivered. This study presents BuildOcc, an open-source Python platform that grounds large language model agents in the American Time Use Survey (ATUS), a nationally representative diary dataset covering 16,684 respondents. Through BuildOcc, each simulated occupant agent can be instantiated with a demographic persona drawn from ATUS population statistics, a memory stream that accumulates and reflects on timestep-level observations, and an activity scheduler that samples empirically from ATUS time-at-activity distributions. The platform exposes a three-layer interface - Python library, REST API, and Model Context Protocol server - so that any building energy tool (EnergyPlus, Home Assistant) can integrate behavioral intelligence without bespoke coupling code. A plugin registry lets the community add new occupant strata, custom schedulers, and alternative memory backends as separate installable packages. Two validation tiers show that ATUS-grounded sampling reproduces empirically calibrated activity distributions and that demographic priors propagate into persona-consistent agent reasoning across timesteps, establishing internal consistency across strata. BuildOcc provides the building energy community with a reusable, openly available implementation of the occupant behavioral layer. BuildOcc is openly released at https://doi.org/10.5281/zenodo.21192895 under the Apache License 2.0 and installable via pip install buildocc.","authors":["Wooyoung Jung"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-28","first_seen":"2026-09-03","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.02729","pdf_url":"https://arxiv.org/pdf/2609.02729","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","建筑能源","ATUS数据"],"reason":"用LLM agent模拟建筑内人员行为，基于ATUS真实数据对照，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":5,"question":"如何构建一个以美国时间使用调查（ATUS）为数据基础、基于大语言模型（LLM）的居住者智能体平台，用于建筑能耗研究中的居住者行为仿真？","design":"该研究开发了 BuildOcc 平台，使用 LLM 智能体模拟四类美国人口群体（在职单身成年人、退休夫妇、在职父母、无业成年人）。每个智能体包含从 ATUS 人口统计中抽取的人口特征、从 ATUS 数据中采样的活动日程、记忆流和推理引擎。在每个 15 分钟时间步，智能体根据人口特征、记忆和当前环境选择下一个活动并说明理由，同时也会对需求响应信号做出接受、拒绝或推迟的决定。平台提供 Python 库、REST API 和 MCP 服务器三层接口，便于与 EnergyPlus 等工具集成。","baseline":"美国时间使用调查（ATUS），包含 16,684 名受访者的全国代表性日记数据，其中 6,611 名受访者属于四个目标人口阶层。","findings":"第一层验证表明，活动调度器能够将 ATUS 活动分布复现到采样噪声范围内；第二层验证表明，人口先验信息能够传播到智能体的推理中，产生与人口特征一致的行为差异。","reliability":"论文承认的局限包括：每个时间步只能选择一个活动；ATUS 仅覆盖美国人口；活动仅基于主要活动（ATUS 一次只记录一个活动）；记忆重要性分数由智能体自行分配并用于检索，缺乏外部校准或反馈路径。","relevance":"该研究将 LLM 智能体与全国代表性时间使用调查数据结合，用于模拟人类行为，并进行了与真实数据的对照验证，属于人类仿真研究，但场景限定于建筑能耗领域，与经济金融问题关联度较低。","inspiration":"借鉴其将 LLM 智能体与大规模调查数据结合、通过分层抽样和记忆机制生成个体行为的方法，可用于构建具有人口代表性的经济决策仿真。｜可迁移到消费者跨期选择或家庭能源消费行为研究，例如模拟不同人口群体对动态电价或节能政策的反应。｜设计雏形：以美国消费者支出调查（CEX）或收入动态面板研究（PSID）为数据基础，构建 LLM 智能体代表不同收入阶层，施加电价上涨或补贴政策处理，结果变量为能源消费和支出变化，并与真实调查数据对照验证。"}},{"id":"2609.30940","version":1,"title":"Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms","zh_title":"LLM智能体社会中的金融脆弱性：协调失败与稳定机制","abstract":"Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\\% of baseline bank-run episodes and 83\\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.","authors":["Zhenhao Fu","Ruipeng Xu","Qibing Ren"],"categories":["cs.AI","q-fin.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30940","pdf_url":"https://arxiv.org/pdf/2609.30940","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","金融仿真","协调失败"],"reason":"用LLM agent模拟金融系统中的协调失败，涉及经济场景，但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":6,"question":"多个LLM智能体在共享金融环境中是否会自发产生系统性金融脆弱性，以及如何通过交互机制设计来缓解这种脆弱性？","design":"使用FRAIL框架，将七个主流LLM作为金融决策者，分别置于银行挤兑、债务展期和奖励众筹三种动态金融环境中，通过多轮交互模拟决策，测量系统失败率（如银行倒闭、债务违约、众筹失败）和承诺形成的时间模式。","baseline":"无对照","findings":"在基线条件下，77%的银行挤兑和83%的债务展期情景以失败告终，即使没有智能体被指示破坏系统。三种交互机制（补偿承诺、集中承诺协议、参与者主导联盟）均能改善结果，但最佳机制因金融结构而异，成功稳定依赖于在防御行为自我强化之前尽早形成广泛承诺。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟金融协调失败，属于经济场景下的多智能体仿真，但缺乏真实人类数据对照，适合关注LLM仿真方法本身或金融AI安全的研究者阅读。","inspiration":"借鉴其动态金融环境设计和多智能体交互机制比较，可迁移到资产定价实验或政策公告预期形成等场景。｜可设计一个信贷审批歧视实验，用LLM扮演银行信贷员和借款人，施加不同信息披露政策作为处理，测量贷款批准率和违约率，并与真实信贷数据对照。"}},{"id":"2609.31054","version":1,"title":"Cheap, open agents make LLM pollution harder to mitigate","zh_title":"廉价开放智能体使LLM污染更难缓解","abstract":"Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.","authors":["Raluca Rilla","Anne-Marie Nussberger","Rui Mata","Dirk U. Wulff"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31054","pdf_url":"https://arxiv.org/pdf/2609.31054","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM污染","调查数据","检测方法"],"reason":"研究LLM污染人类调查数据，评估检测方法，与仿真可靠性直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":7,"question":"完全开源的自主智能体是否降低了LLM污染人类调查数据的门槛，以及现有检测方法能否有效识别这些智能体？","design":"使用九种智能体配置（三种完全开源、两种混合、四种专有）自主完成同一份调查问卷，每种配置运行40次，共360次运行。智能体被赋予随机的人口统计特征（性别、年龄），并收到统一指令。调查包含多种回答类型和18项检测检查（16项通过/失败检查和2项计时测量），用于评估智能体的表现和可检测性。","baseline":"3,242名来自Prolific的美国参与者，在2025年10月20日至11月4日期间完成同一调查问卷，样本在性别、年龄和种族上大致代表美国人口。","findings":"完全开源的智能体在本地运行且无使用费用，其表现与商业替代方案相当。开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，但开放式文本回答在区分智能体和人类方面效果最好。","reliability":"论文指出，开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，因此需要多层检测策略，特别是强调开放式文本分析。此外，研究仅操纵了年龄和性别两个人口统计变量，未涵盖其他可能影响回答的变量；智能体被明确指示避免与可疑嵌入命令交互，这可能降低了某些检测的失败率。","relevance":"该研究直接评估了LLM智能体对人类调查数据的污染风险及检测方法，与您关注的LLM仿真可靠性和偏差问题高度相关，特别是它提供了真实人类数据作为对照，并揭示了开源智能体带来的新风险，值得精读原文以了解具体检测方法和失效模式。","inspiration":"借鉴其多层检测策略和开放式文本分析来识别LLM生成回答的方法，可迁移到经济金融领域的调查数据质量控制中。｜可应用于消费者信心调查、投资者情绪调查或政策评估中的问卷数据，检测是否存在LLM污染。｜设计：以真实人类调查数据（如密歇根大学消费者信心调查）为基准，让开源和商业LLM智能体自主完成同一问卷，比较其回答分布和开放式文本特征，并开发基于文本分析的检测指标，评估不同检测方法的敏感性和特异性。"}},{"id":"2608.27167","version":2,"title":"Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable","zh_title":"校准到知道，但未校准到行动：伪造证据使LLM智能体对不可知问题做出承诺","abstract":"An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.","authors":["Pranav Aggarwal"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-08-28","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2608.27167","pdf_url":"https://arxiv.org/pdf/2608.27167","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM决策偏差","可靠性评估","批判性研究"],"reason":"研究LLM在不可知问题上的决策偏差，评估其可靠性，批判性指出失效条件，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":8,"question":"在不可知（aleatoric）问题上，专业外观的证据面板是否仅凭其权威包装就能诱导LLM智能体做出方向性承诺，而非真正提供信息？","design":"研究用12个前沿LLM作为被试，在四个领域（股票、加密货币、体育、天气）构造不可知问题，施加不同证据梯度（无证据、薄面板、丰富面板、乱序面板、全伪造面板），测量模型是否做出方向性承诺（行动）以及陈述的概率，并与可回答的匹配问题对照。","baseline":"无对照","findings":"伪造证据面板与真实面板诱导的承诺率在统计上无差异，表明触发行动的是包装的权威性而非信息内容；模型的陈述概率几乎不随证据梯度变化，且比气候学基线更差，但同一模型在可回答问题上几乎完美作答。","reliability":"论文承认天气领域无密封结果，面板提供的集合预报概率可能具有真实技能，因此该领域作为不可知性工具最弱；训练出的门控在刚性响应格式下失效，需要留出推理空间。","relevance":"该研究直接评估LLM在不可知问题上的决策偏差，揭示仿真失效的特定条件（证据包装触发行动），对关注LLM仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其通过伪造证据与真实证据对比来分离信息与包装的因果设计，以及用密封结果验证不可知性并测量行动而非仅测信念的做法｜可迁移到资产定价实验或政策公告的预期形成研究，例如测试LLM代理在呈现专业外观的虚假市场数据时是否会产生过度自信的交易决策｜用LLM作为被试，随机分配真实与伪造的市场分析面板，测量其买卖决策和置信度，并与人类实验数据或历史市场结果对照，检验权威包装对决策的影响。"}},{"id":"2609.30867","version":1,"title":"Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations","zh_title":"气候政策因果评估中识别假设的证据基础审计","abstract":"Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS","authors":["Yonghong Zhang","Yong Xie","Isabel M. Parra","Ricardo Correia"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30867","pdf_url":"https://arxiv.org/pdf/2609.30867","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A4","B3"],"tags":["LLM审计","因果推断","方法论"],"reason":"提出LLM审计因果推断假设的方法论框架，可迁移到仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":10,"question":"如何用大语言模型审计气候政策因果评估中双重差分法识别假设的证据支持度？","design":"本文提出 ARGUS 流水线，用大语言模型对双重差分研究的 11 个识别维度进行证据检索与充分性评估，并输出风险报告；模型仅负责检索和评估，控制流固定。","baseline":"无对照","findings":"在注入缺陷基准上，ARGUS 检出 73% 的植入缺陷，而关键词流水线仅检出 18%；在 26 篇经济学论文上，约 40% 的论文-维度评估因缺乏可检索证据而弃权。","reliability":"论文承认在真实论文上证据覆盖不足导致弃权率高，且过度严厉导致与人工标签一致性低；气候政策专用语料上的表现未测试。","relevance":"该研究提出用 LLM 审计因果推断假设的方法论框架，可迁移到仿真可靠性评估，值得阅读原文了解其证据门控与弃权机制。","inspiration":"值得借鉴的是将识别假设分解为可审计的维度，并用检索门控控制 LLM 的评估范围，避免模型自由发挥｜可迁移到政策评估中的因果推断可靠性审计，例如碳税、补贴或监管政策的效果评估｜可设计一个仿真实验：用 LLM 扮演审稿人，对一组已发表的经济学论文的识别假设进行审计，处理是提供不同质量的证据，结果变量是风险评级，并与人类专家评级对照。"}},{"id":"2609.31245","version":1,"title":"RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models","zh_title":"RupeeBias：审计印度经济指导中大语言模型的人口统计偏差","abstract":"Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.","authors":["Pavithra P M Nair","Bhavik Talaviya","Shourya Bhushan","Rahul Pankajakshan","Seema Guruvadoo","Avinash Agarwal","Gilad Gressel","Krishnashree Achuthan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31245","pdf_url":"https://arxiv.org/pdf/2609.31245","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM偏差审计","经济决策仿真","人口统计代表性"],"reason":"用LLM生成经济建议并审计人口统计偏差，有真实人类数据对照，涉及经济决策场景，…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":13,"question":"在印度经济咨询场景中，大语言模型生成的经济建议是否因用户的人口统计特征（如种姓、宗教、性别等）而产生系统性偏差？","design":"构建 RupeeBias 基准，包含 39,150 个提示，覆盖薪资估算、加薪估算、还价建议和服务定价建议四个用例。采用单属性反事实设计，固定用户资质、经验或服务描述，仅改变一个印度特定的人口统计标识符（共 87 个，涵盖种姓、宗教、地域、性别、残疾和城乡位置六个轴），并同时用英语和印地英语构建提示。评估九个 LLM 在这些提示上的输出差异。","baseline":"无对照","findings":"在仅人口统计标识符不同的相同提示下，LLM 生成的经济输出平均差异达 20.2%，且在所有六个轴上都存在系统性人口统计差异。","reliability":"论文未讨论","relevance":"该研究直接针对 LLM 在经济学场景中的仿真偏差，使用反事实设计审计人口统计特征对经济建议的影响，与研究者关注的 LLM 人类仿真可靠性及偏差问题高度相关，值得阅读原文了解具体偏差模式和测量方法。","inspiration":"借鉴其单属性反事实设计，通过仅改变一个属性来隔离人口统计特征对模型输出的因果影响，并构建大规模、多语言、多场景的提示集｜可迁移到信贷审批歧视、保险定价、工资谈判等经济金融决策场景，审计 LLM 在金融建议中的群体偏差｜以 LLM 作为虚拟被试，处理为在贷款申请或保险报价提示中改变申请人姓名、性别或种族等标识，结果变量为模型给出的利率、保费或授信额度，并与真实信贷或保险数据中的群体差异进行对照，检验模型偏差是否反映或放大现实歧视。"}},{"id":"2609.31013","version":1,"title":"Same Text, Different Numbers: The Divergence of LLM-Based Measures","zh_title":"相同文本，不同数字：基于LLM的测量分歧","abstract":"Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.","authors":["Hamid Boustanifar","Sasan Mansouri"],"categories":["cs.AI","cs.CL","q-fin.GN","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31013","pdf_url":"https://arxiv.org/pdf/2609.31013","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["LLM测量一致性","算法保真度","文本分析"],"reason":"评估LLM文本测量跨模型一致性，涉及测量偏差与统计推断，有真实数据对照，方法可…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":11,"question":"LLM 生成的文本测量在多大程度上不受模型选择影响？","design":"用七家不同提供商的 LLM 对同一批标普 500 公司财报电话会议问答环节文本，按相同提示词对 13 个构念打分，比较跨模型分数的一致性、方差来源及对下游回归推断的影响。","baseline":"以 Loughran-McDonald 等既有词典法文本测量作为对照，并考察 LLM 分歧与分析师预测分歧、市场反应等真实人类判断的关联。","findings":"跨模型平均秩相关仅 0.52，且构念间差异大；方差分解显示转录本共同成分平均只占 34%，模型特定成分显著。模型选择会实质改变下游回归系数的符号、大小和显著性，而 LLM 分歧与后续分析师或市场分歧无关。","reliability":"论文指出 LLM 测量应视为模型依赖的，需跨提供商验证；自报置信度不能识别更可靠的观测，规则化提示词反而可能增大分歧，且大部分分歧无法由可观测特征解释。","relevance":"直接评估 LLM 作为测量工具在金融文本分析中的可靠性，揭示模型选择对实证结论的威胁，对关注仿真偏差和稳健性的研究者有重要参考价值。","inspiration":"借鉴其多模型同文本打分并做方差分解的设计，可系统检验测量工具的选择效应｜可迁移到政策公告文本的情绪或不确定性测量、信贷审批中的软信息编码、分析师报告语调等场景｜用多家 LLM 对同一批政策声明或贷款申请文本按相同提示词打分，以人工编码或市场反应数据为基准，比较跨模型一致性及对后续回归推断的影响。"}},{"id":"2609.31468","version":1,"title":"PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents","zh_title":"PriceBench：LLM预订代理中价格、质量与品牌偏好的诊断基准","abstract":"LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \\$247 to \\$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.","authors":["Pavel Kireyev"],"categories":["econ.GN","cs.AI","cs.CL","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31468","pdf_url":"https://arxiv.org/pdf/2609.31468","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","消费者选择","偏好测量"],"reason":"用LLM替代人类消费者进行预订选择，并与真实酒店数据对照，属于经济学场景仿真","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":14,"question":"LLM作为酒店预订代理时，其价格、质量和品牌偏好是什么，这些偏好在不同LLM之间有何差异？","design":"构建PriceBench基准，包含3600个酒店选择任务（1800个二元、1800个三元），基于179家真实纽约酒店，属性包括星级、价格、评分等。对28个LLM进行测试，每个任务以原始顺序和交换顺序各评分一次，使用logit选择模型从选择中恢复偏好参数。","baseline":"无对照（论文未使用真实人类预订数据作为基准，但引用了人类预订者对质量的估值范围进行对比）。","findings":"能力与选择的一致性相关，而非选择内容：更强的LLM偏好更一致，较弱的LLM要么锁定位置（受列表顺序控制），要么几乎无差异。价格敏感度跨LLM差异超过一个数量级，价格/质量权衡导致相同任务下平均每晚预订价格从247美元到393美元不等。","reliability":"论文讨论了位置锁定LLM在温度0下失效，采样无法恢复偏好；偏好幅度受提示格式、解码方式和选项数量影响；品牌偏好与提供商无关。","relevance":"该研究用LLM替代人类消费者进行预订选择，属于经济学场景仿真，但缺乏真实人类对照，主要提供LLM间差异的诊断，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其通过随机化属性、重复呈现和logit模型识别偏好的方法，可迁移到消费者选择、定价策略等经济问题。｜例如，在信贷审批歧视研究中，用LLM扮演信贷员，改变申请人属性（如种族、性别）观察决策差异。｜设计：以LLM为被试，呈现贷款申请（处理为不同种族/性别信号），结果变量为批准决策，对照真实信贷员历史数据评估偏差。"}},{"id":"2609.30705","version":1,"title":"The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?","zh_title":"思考的代价：测试时推理在LLM交易中是否值得？","abstract":"While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.","authors":["Jiayi Chen","Guiling Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30705","pdf_url":"https://arxiv.org/pdf/2609.30705","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM交易","推理成本","经济决策"],"reason":"用LLM模拟金融决策并与真实市场数据对照，涉及经济场景和失效条件，但非人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":9,"question":"在固定模型、信息、提示、输出格式和投资组合规则的情况下，增加推理努力是否可靠地改善LLM交易组合的净收益？","design":"对DeepSeek、GPT和Gemini三个模型家族的代表性LLM进行受控干预，改变推理努力水平（无、低、高、最大），在2024年241个形成日对100只美国流动性股票进行评分，基于数值特征、可识别新闻和掩蔽新闻三种输入条件，构建等权市场中性投资组合（买入前10、卖空后10），持有5天，扣除10个基点的交易成本，比较不同推理水平下的净收益。","baseline":"无对照","findings":"在三个模型家族中，增加推理并未可靠地改善净投资组合收益；DeepSeek从无推理到最大推理的表现是非单调的。重复生成显示处理效应和投资组合选择不稳定，即使总体得分相似。","reliability":"论文承认推理水平的变化会改变输出结构可靠性，但可靠性提高并不自动意味着收益提高；重复生成导致处理效应不稳定，表明结果可能因随机生成而异；研究仅覆盖一年美国股票数据，且置信区间较宽，不能证明所有效应为零。","relevance":"该研究用LLM模拟金融决策并与真实市场数据对照，评估推理努力对经济结果的影响，属于经济场景下的仿真研究，并揭示了仿真失效条件（推理增加不带来稳定收益），值得阅读原文以了解其受控设计和稳健性检验方法。","inspiration":"借鉴其受控干预设计：在固定其他因素下仅改变推理努力，并使用重复生成和冻结审计评估稳定性｜可迁移到资产定价实验或投资决策仿真，检验LLM推理深度对预测准确性和组合表现的影响｜以LLM为被试，处理为不同推理水平，结果变量为组合净收益或预测误差，用真实历史市场数据作为基准对照。"}},{"id":"2609.31095","version":1,"title":"Confident, Not Wiser: The Dunning-Kruger Effect in Human-AI Interaction","zh_title":"自信而非更明智：人机交互中的达克效应","abstract":"AI assistance can improve performance without improving self-assessment. We report a study (N=366) comparing Human alone and Human+AI performance on reasoning tasks, for which the AI model is benchmarked on the same items. Participants estimated global and block performance and rated confidence in their answers. Human+AI achieved higher scores, but self-estimates tracked performance weakly. Average overestimation was similar across groups, covering individual errors. Across tasks, confidence distinguished correct from incorrect answers less accurately in the Human+AI group, while within-task differences remained uncertain. The Dunning-Kruger pattern was found in both groups, with a larger observed contrast in Human+AI. Controls for score noise reduced but did not eliminate the pattern, with the controlled group difference remaining inconclusive. An extended computational account describes global and block estimates. Our findings distinguish performance augmentation from metacognitive augmentation and motivate interfaces that support verification, communicate task-specific AI model performance, and help users evaluate the quality of their joint work rather than produce answers.","authors":["Daniela Fernandes","Michelle Rausch","Agnes Mercedes Kloft","Daniel Buschek","Robin Welsch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31095","pdf_url":"https://arxiv.org/pdf/2609.31095","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["人机交互","元认知","AI辅助决策"],"reason":"研究人类与AI协作中的元认知偏差，有真实人类数据对照，结论可迁移到LLM仿真可…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":12,"question":"人类与AI协作时，AI辅助是否在提升任务表现的同时改善元认知（自我评估准确性、信心区分度），以及邓宁-克鲁格效应在有无AI辅助下如何表现。","design":"本研究并非LLM仿真研究，而是人类实验：366名被试分为Human alone和Human+AI两组，完成40道推理题（矩阵、心理旋转、三段论、字母串），AI为GPT-5.6 Luna，并单独对AI模型进行基准测试。被试在答题前后进行全局和分块表现估计，并对每题答案给出信心评分。","baseline":"Human alone组作为人类基准，同时AI模型单独在相同题目上的表现作为AI基准。","findings":"Human+AI组得分更高，但自我估计与实际表现的相关性弱，平均高估程度与Human alone组相似；信心区分正确与错误答案的能力在Human+AI组更差。邓宁-克鲁格模式在两组均存在，Human+AI组观察到的对比更大，但控制分数噪声后组间差异不明确。","reliability":"论文未明确讨论失效条件，但指出研究为探索性观察研究，未预注册；控制分数噪声后邓宁-克鲁格效应的组间差异不明确，且任务内信心区分度差异不确定。","relevance":"该研究直接涉及人类与AI协作中的元认知偏差，有真实人类数据对照，结论可迁移到LLM仿真中关于自我评估和信心校准的建模，值得阅读原文以了解具体测量和稳健性检验方法。","inspiration":"借鉴其同时测量全局估计、分块估计和逐题信心，并设置Human alone和Human+AI对照以及AI单独基准，以分离表现提升与元认知变化。｜可迁移到金融决策场景，如投资者使用AI辅助进行资产配置或信贷审批中，评估AI辅助是否改善决策质量但未改善决策者的信心校准。｜设计实验：招募投资者作为被试，随机分为仅人类组和人类+AI组，完成一系列投资决策任务，AI为LLM提供建议，结果变量为投资组合收益和投资者对自身表现的估计及信心评分，对照真实市场数据或历史基准。"}},{"id":"2609.31046","version":1,"title":"Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference","zh_title":"用大语言模型和知识图谱引导推理建模学生意义建构","abstract":"Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. We investigate whether instruction-tuned large language models (LLMs) can support multidimensional analysis of collaborative sensemaking without task-specific training, and whether structured knowledge-state information improves model inference. We evaluate two mid-size LLMs on 23 richly annotated, expert-labeled episodes across prompting conditions that vary definitional scaffolding, reasoning mode, and turn structure. Without reasoning, models tend to overpredict successful sensemaking; reasoning-enabled prompting improves identification of unsuccessful cases. Knowledge-state diagnostics provide additional grounding, improving detection of unsuccessful sensemaking and increasing agreement with expert annotations. No single configuration performs best across all sensemaking dimensions, underscoring the multidimensional nature of the task.","authors":["\\\"Ozge Alacam","Z\\\"ubeyde Demet Kirbulut G\\\"une\\c{s}","Funda Ekici","Nurcan Turan-Oluk","Dilay Din\\c{c}demir","Hakk{\\i} Kaday{\\i}f\\c{c}{\\i}","Sevin\\c{c} Nihal Ye\\c{s}ilo\\u{g}lu","Burcu I\\c{s}{\\i}k","Halil T\\\"umay","Sinem Gencer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31046","pdf_url":"https://arxiv.org/pdf/2609.31046","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","教育对话分析","知识图谱"],"reason":"LLM用于分析学生对话，替代人工标注，而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.31078","version":1,"title":"OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas","zh_title":"OmouAI：基于模拟人格的论证式人机政策协商","abstract":"Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in combination with computational argumentation, a field which excels in representing and reasoning within debates. OmouAI allows a human user to deliberate policy claims for real-world challenges with simulated personas, e.g., representing stakeholders, domain experts or devil's advocates, towards reducing sycophancy. Each persona generates its own arguments, and the arguments of all parties form a shared argumentation framework. Users can then contest, add and revise arguments, providing crucial human oversight. Then, arguments are evaluated using deterministic argumentative semantics against external goals, such as the UN Sustainable Development Goals, guaranteeing faithful explanations. The advancement or worsening of the goals thus serve as indicators for the policy recommendations.","authors":["Stylianos Loukas Vasileiou","Antonio Rago","William Yeoh","Georgina Curto"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31078","pdf_url":"https://arxiv.org/pdf/2609.31078","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","政策协商","计算论证"],"reason":"用LLM模拟persona进行政策辩论，属社会过程模拟，但无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.31607","version":1,"title":"Statistical attribute alignment for black-box generative AI via output post-processing","zh_title":"通过输出后处理实现黑盒生成式AI的统计属性对齐","abstract":"Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \\rightarrow \\infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.","authors":["Kevin Jiang","Morgane Austern","Edgar Dobriban","Jason M. Klusowski"],"categories":["stat.ME","cs.AI","cs.LG","math.ST","stat.TH"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31607","pdf_url":"https://arxiv.org/pdf/2609.31607","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["生成式AI","属性对齐","合成数据"],"reason":"涉及生成合成数据对齐目标分布，但非仿真人类被试，而是后处理调整属性分布。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.30716","version":1,"title":"Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4","zh_title":"言语胜于顺序：Gemma 4 的行为评估","abstract":"When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b model across a targeted behavioral suite (n = 13 items, 784 forward passes in short, single-turn contexts) using a completely counterbalanced experimental design. This setup allowed us to mathematically isolate the specific effects of source framing and reading position, while ensuring the model's natural vocabulary biases were canceled out. Across ten test conditions, we discovered the following: 1. Source framing heavily overpowers reading position. When directly competing, the semantic framing of a source (such as presenting it as an official guideline or a fresh update) had a significantly stronger impact on the model's final answer than the presentation order of the document. 2. The model favors the first document it reads, but this bias is highly variable. While the model consistently demonstrated a primacy effect (preferring the first document presented), the actual strength of this bias fluctuated by at least a factor of 5 based solely on the surface wording. 3. Overall structural repetition, not short copy-cues, drives positional bias. The model's preference for the first document is not a mechanical reaction to short, repetitive trigger phrases, such as \"is [Answer]\". However, the primacy effect does increase significantly when the two competing documents are structurally identical, using word-for-word verbatim templates. Introducing variation in the overall wording between the two sources reduces this positional bias.","authors":["Amanda Fitch"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30716","pdf_url":"https://arxiv.org/pdf/2609.30716","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为评估","来源框架","顺序效应"],"reason":"评估模型行为而非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.30986","version":1,"title":"Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries","zh_title":"评估中文大语言模型在基于在线搜索查询的事实问题上的谄媚行为","abstract":"As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or causes uncertain responses to become belief-aligned incorrect answers. It also remains unclear whether anti-sycophancy interventions preserve or restore factual accuracy or merely shift responses toward uncertainty. We analyze factual sycophancy in Chinese-language information seeking using yes/no fact-checking questions. Our analysis covers 364,941 responses from three frontier Chinese-based LLMs (DeepSeek, Qwen, and Doubao) to 12,165 factual questions derived from real-world Chinese search queries. We evaluate the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses. Under incorrect user beliefs, we distinguish belief-aligned errors from losses of factual confidence, in which initially correct answers become uncertain. Patterns vary across models and reasoning settings: reasoning is not a consistent safeguard, and anti-sycophancy instructions can reduce incorrect agreement while increasing uncertainty. In Chinese-language factual question answering, avoiding agreement with false beliefs is therefore not equivalent to preserving factual accuracy, highlighting the value of transition-level evaluation. Such behavior may undermine the reliability of LLM-mediated information access by reinforcing misinformation or weakening users' confidence in factually correct answers.","authors":["Geng Liu","Feng Li","Mengxiao Zhu","Francesco Pierri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30986","pdf_url":"https://arxiv.org/pdf/2609.30986","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM谄媚","事实准确性","中文模型"],"reason":"研究LLM在事实问答中的谄媚行为，测量模型本身而非仿真人类被试，无人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.31506","version":1,"title":"Evaluating Cultural Awareness of LLMs for Haitian Creole","zh_title":"评估LLM对海地克里奥尔语的文化意识","abstract":"Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring portrayals of Haitian characters through hardship and resilience, showing that even positive characterizations can encode stereotypical narratives. Our code, benchmark, and evaluation framework are publicly available.","authors":["Christelle Clervilsson","Yanzhu Guo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31506","pdf_url":"https://arxiv.org/pdf/2609.31506","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["文化意识","低资源语言","模型评估"],"reason":"评估LLM的文化意识，测量模型而非仿真人类被试，但涉及文化偏差，与仿真可靠性相…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.31603","version":1,"title":"User Model Extraction via Belief Self-Distillation","zh_title":"通过信念自蒸馏提取用户模型","abstract":"Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.","authors":["Ali Holmov","Yiran Huang","Kirill Bykov","Zeynep Akata"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31603","pdf_url":"https://arxiv.org/pdf/2609.31603","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM内部表征","用户建模","可解释性"],"reason":"研究LLM对用户的隐式建模，属于模型内部状态测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.31260","version":1,"title":"Agentic Limit Order Books: Phase Transitions and Market Impact","zh_title":"智能体限价订单簿：相变与市场冲击","abstract":"We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fundamental quantitative phenomena: equilibrium phase transitions in order flow regime shifts, and the structural dynamics of market impact. We show that agentic LOBs exhibit distinct phase boundaries separating orderly price discovery from hyper-volatile cascade states, governed by critical thresholds in the number of agents and observable market depth. Furthermore, we demonstrate that market impact under agentic liquidity provision deviates from classical square-root dynamics, exhibiting distinct dissipative, balanced, and non-dissipative regimes under non-linear feedback loops.","authors":["Jan Rosenzweig"],"categories":["q-fin.TR","cs.AI","q-fin.CP"],"primary_category":"q-fin.TR","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31260","pdf_url":"https://arxiv.org/pdf/2609.31260","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体市场模拟","限价订单簿","强化学习"],"reason":"用RL agent模拟限价订单簿市场动态，属社会/经济过程模拟，但无真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2608.29420","version":2,"title":"One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI","zh_title":"一种能力还是多种？基准效度的结构与预测检验对前沿AI经济基准意见不一","abstract":"Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct validity that a structural test and a predictive test can answer in opposite ways. We show that they do on a hash-pinned snapshot of a frontier leaderboard with 421 model configurations across twelve benchmarks, four of them economic, of which 103 configurations carry all three sparsely scored economic benchmarks and 96 carry all twelve; four hypotheses and their thresholds were fixed before analysis, and every deviation from the plan is reported. The first factor of a three-factor extraction carries 74.5% of common variance and tracks release date (R^2 = 0.505), and date adjustment lowers its share by 14.9 points. Under the dimensionality rule fixed in advance the economic benchmarks form no factor of their own. A leave-one-benchmark-out test with factors re-estimated inside every fold nevertheless finds that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]) under the linear learners that fit best, an advantage that reverses for tree learners. Under the linear learners the same representation also predicts the eight other benchmarks better, so the battery carries predictive structure that one index misses and the economic benchmarks share it without forming a distinct factor. Construct validity should therefore be assessed by predictive tests alongside structural ones. We give a two-test protocol for benchmark builders and release the pinned data, the analysis plan and the code.","authors":["Louis Yiven Zhu"],"categories":["cs.LG","cs.CY","cs.SE"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-09-01","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2608.29420","pdf_url":"https://arxiv.org/pdf/2608.29420","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准评估","构念效度","AI评测"],"reason":"论文评估经济基准的构念效度，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28504","version":2,"title":"AI in Science: Early Insights","zh_title":"AI在科学中的早期洞察","abstract":"Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements-- LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.","authors":["Mihai Codreanu","Alex Imas","Juan Mateos-Garcia","Joseph Emmens","Evalyne Muiruri","Arthur Turrell","Julian Jacobs","Atoosa Kasirzadeh","Ana Trisovic","Yiyuan Chen","Tanya Rodchenko","Catherine Pollard","Scott Strand","Daniel Rock","Zanna Iscenko","Fabien Curto Millet","Neil Thompson","James Manyika"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-09-25","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.28504","pdf_url":"https://arxiv.org/pdf/2609.28504","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI与科学","生产力","科学过程"],"reason":"研究AI对科学的影响，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.30492","version":1,"title":"Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs","zh_title":"打破同质性：为创造性LLM输出多样化角色集","abstract":"Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform-coverage sampling, and evolutionary persona generation. Evaluations on the Alternative Uses Task (AUT), Infinity-Chat, and Divergent Association Task (DAT) show the benefits of the proposed methods across tasks and creativity objectives. On AUT, evolutionary persona generation increases response diversity by 78.8%, originality by 26.1%, flexibility by 49.5%, and holistic creativity by 13.9% over task-only prompting, while maintaining 98.5% validity; on Infinity-Chat, it nearly doubles persona-induced response separation relative to random personas. Moreover, evolutionary personas compose with creativity-optimized prompting, further increasing its response diversity by 18.6% and creativity by 6.3%. These results establish persona-set geometry as a task-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization.","authors":["Sang Bin Moon","Nicole Cho","Daniel Borrajo","Sumitra Ganesh","Abolfazl Hashemi"],"categories":["cs.CL","cs.AI","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30492","pdf_url":"https://arxiv.org/pdf/2609.30492","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色多样化","创造力生成","提示工程"],"reason":"研究通过多样化角色设定提升LLM创造力输出，属于角色扮演与生成多样性，无人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.30558","version":1,"title":"Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms","zh_title":"通过认知实验范式探究智能体记忆中的稳定性-可塑性权衡","abstract":"Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at https://github.com/jq-ding/MemProbe.","authors":["Jiaqi Ding","Guorong Wu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30558","pdf_url":"https://arxiv.org/pdf/2609.30558","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体记忆","认知诊断","多智能体系统"],"reason":"评估智能体记忆系统，不涉及人类行为仿真或对照，属多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.30897","version":1,"title":"From annotation to reasoning: Culture in language models","zh_title":"从标注到推理：语言模型中的文化","abstract":"How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.","authors":["Daniel Hershcovich","Alexander Conroy","Jens Bjerring-Hansen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30897","pdf_url":"https://arxiv.org/pdf/2609.30897","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文化基准","文学解释","模型评测"],"reason":"论文关注文化基准评测与文学解释，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.30768","version":1,"title":"Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More","zh_title":"思考有助于公平吗？推理令牌解决一些偏差但创造更多","abstract":"Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.","authors":["Deng Pan","Joe Germino","Yihong Ma","Elizabeth Daly","Nuno Moniz","Ting Hua","Nitesh Chawla"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30768","pdf_url":"https://arxiv.org/pdf/2609.30768","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["公平性","推理模型","偏差"],"reason":"研究LLM推理对公平性影响，属模型偏差评测，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.30939","version":1,"title":"MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory","zh_title":"MACBT：具有纵向记忆的多智能体认知行为疗法决策支持系统","abstract":"Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, cognitive restructuring, behavioral experiments, and treatment monitoring) into five collaborative agents. CD Memory tracks cognitive-distortion type, frequency, severity, and restructuring efficacy across sessions to generate pre-session pathology reports and intervention-priority recommendations. We construct a Chinese CBT dialogue corpus via dual-role large language model simulation and train a Qwen3-14B backbone with supervised fine-tuning and direct preference optimization. Evaluation with GPT-4 judges shows MACBT outperforms MeChat, SoulChat, PsyChat, and CPsyCounX in professionalism (2.62) and clinical authenticity (2.25). The full memory-augmented system further improves session quality by 12.6% and achieves a longitudinal mean of 2.29 on cross-session continuity, intervention progression, and personalization.","authors":["De Jiang","Shuo Zhang","Weiwei Liao","Jianying Zhang","Chuanhui Yu","Hongen Liao","Kehong Yuan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30939","pdf_url":"https://arxiv.org/pdf/2609.30939","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1","C3"],"tags":["多智能体系统","认知行为疗法","决策支持"],"reason":"多智能体CBT决策支持系统，无人类行为对照，属角色扮演与协作任务，非仿真被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.31184","version":1,"title":"Accounting for Bias Enables Sustainable LLM Evaluation","zh_title":"考虑偏差实现可持续的LLM评估","abstract":"LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.","authors":["Harshita Katoch","David Antony Selby","Gerrit Gro{\\ss}mann","Sebastian Vollmer"],"categories":["cs.AI","cs.LG","stat.AP","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31184","pdf_url":"https://arxiv.org/pdf/2609.31184","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","偏差校正","测量模型"],"reason":"论文研究LLM作为评判者的评估偏差校正，属于纯NLP评测方法，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.31215","version":1,"title":"DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration","zh_title":"DIAL：基于自适应人类偏好校准的位置去偏LLM评判器","abstract":"Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.","authors":["Zesheng Cai","Yingqi Fan","Sichang Chen","Jin-Hong Du"],"categories":["cs.AI","stat.AP","stat.ME","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31215","pdf_url":"https://arxiv.org/pdf/2609.31215","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判","偏好校准","评估方法"],"reason":"论文研究LLM作为评判者的位置偏差与人类偏好校准，属于评估方法改进，非用LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.31473","version":1,"title":"Game Arena: Strategic LLM Evaluation in Competitive Environments","zh_title":"游戏竞技场：竞争环境中的大语言模型战略评估","abstract":"We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.","authors":["Bovard Doerschuk-Tiberi","Yao Yan","Justin Chiu","Hann Wang","Timothy Chung","Martyna Plomecka","John Schultz","Jon Lipovetz","Clayton Drazner","Yuchen Zhuang","Jaimie Hwang","Nate Keating","Riley Jones","Andrew Lee","Oran Kelly","Ian Gemp","Michael Aaron","Laurel Prince","Kate Larson","Jeff Moser","Harrison Jobe","Chad Woodford","Siqi Liu","Andrew Wang","Bo Chang","Christopher D'Mello","Diane Chaleff","Addison Howard","Johnny Yip","Chuck Sugnet","Antonio Gulli","Meghan O'Connell","Will Cukierski","Nenad Tomasev","Dima Yeroshenko","Kinjal Parekh","Roxanne Daniel","Marc Lanctot","Domino Weir","Elsa Dong","Daniel Hennes","Melissa Nalubwama","Robert Fraser","Ryan Trostle","Jun Peng","Tom Mason","Lloyd Hightower","Chiamaka Chukwuka","Yuexiang Zhai","Phoebe Kirk","Yi Su","Yuting Han","Jie Ren","Chris Prichard","Sahand Sharifzadeh","Karim Hakimzadeh","DJ Sterling","Meg Risdal","Kate Olszewska","Ya Xu","Orhan Firat","Minmin Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31473","pdf_url":"https://arxiv.org/pdf/2609.31473","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评估","游戏AI","多智能体"],"reason":"评估LLM在游戏中的策略能力，属游戏仿真环境，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.31563","version":1,"title":"Multi-agent Scaling Across Disjunctive and Compensatory Tasks","zh_title":"分离型与补偿型任务中的多智能体扩展","abstract":"Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.","authors":["Carolina Fortuna","Blaz Bertalanic"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31563","pdf_url":"https://arxiv.org/pdf/2609.31563","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","任务扩展","LLM协作"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.30466","version":1,"title":"A Benchmarking Framework for Context-aware XR Interfaces","zh_title":"面向情境感知XR界面的基准测试框架","abstract":"Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application as a connected graph of functional facets, each a semantically coherent group of related capabilities that together support a shared user intent. On this representation, we build MineXR++, a dataset augmenting prior XR interface data with facet-level annotations, and formulate three canonical tasks of context-aware suggestion: context factor analysis, initial facet suggestion, and next facet suggestion. Our evaluation protocol scores suggestion methods by a simulated interaction metric, the navigation and search cost of reaching the desired functionality. Through experiments benchmarking global popularity, relational retrieval, and LLM-based methods, we demonstrate that ContextXR enables the systematic, reproducible evaluation of context-aware XR interfaces.","authors":["Hyunsung Cho","Sarah Yewon Yun","Nancy Ruonan Sun","Ben Lafreniere","Mark Parent","Kashyap Todi","Tanya R. Jonker","Hrvoje Benko","Sherry Tongshuang Wu","David Lindlbauer"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30466","pdf_url":"https://arxiv.org/pdf/2609.30466","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["XR界面","基准测试","LLM方法"],"reason":"论文聚焦XR界面基准测试，LLM仅作为方法之一，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.31219","version":1,"title":"Research with AI Agents: How Agentic Systems Are Changing Scientific Work","zh_title":"AI代理研究：代理系统如何改变科学工作","abstract":"Background. Agentic AI systems independently decompose tasks such as literature search, data analysis, and programming into subtasks, search the web, access databases, and execute code. This allows them to perform digital research tasks at high speed. Objectives. Under what conditions does the use of agentic systems produce reliable efficiency gains, and which tasks remain with researchers? Materials and methods. Summary of current studies on literature searches, data analysis, software development, and clinical decision support. Results. For digital activities, work shifts from execution to steering and review. Efficiency gains are greatest when expected behavior can be formalized in advance and tested automatically. In complex agentic systems, recorded sequences of reasoning steps and tool calls can quickly become too extensive for human review. Furthermore, explanations generated by the model do not reliably reflect how an output was produced. One possible step toward more reliable systems is the validation of individual components. The limited reviewability extends beyond research itself; the review of scientific articles and grant proposals is also reaching capacity limits. In pathology, curated and annotated data, researchers' own analytical skills, and institutional exchange of experience are becoming increasingly important. Conclusions. Researchers remain responsible for their results. They must determine what to delegate and how to review the results. The importance of a research question to patients, the field, and society cannot be fully assessed using formalized criteria and remains a matter of expert judgment. Agentic systems can free up time for this.","authors":["Johannes Lotz","Markus Wenzel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31219","pdf_url":"https://arxiv.org/pdf/2609.31219","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","科研自动化","人机协作"],"reason":"讨论AI代理在科研任务中的自动化，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.30427","version":1,"title":"Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation","zh_title":"假新闻理论：利用跨学科洞见进行计算建模、检测与解释","abstract":"Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cross-disciplinary theories of fake news into measurable features for automated detection and explanation through statistical techniques and large language models. To that end, we conduct a structured cross-disciplinary review of theories from social sciences, psychology, economics, among other disciplines that reveal how fake news persuades and spreads, thereby establishing a broad theoretical foundation for computational modeling. Experiments on benchmark datasets show that theory-derived features are predictive and provide interpretable, theory-referenced diagnostic signals. Multi-feature models generally outperform individual features, although gains among the strongest small feature combinations are modest. Our work highlights the value of interdisciplinary perspectives in building robust and interpretable fake news detection systems, advancing the foundation for human-centered approaches in combating disinformation.","authors":["Zhaoyang Cao","Miriam Metzger","Reza Zafarani"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30427","pdf_url":"https://arxiv.org/pdf/2609.30427","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["假新闻检测","可解释性","跨学科理论"],"reason":"论文聚焦假新闻检测与解释，使用LLM提取特征，但非以人类为参照系仿真被试，属纯…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.31590","version":1,"title":"AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs","zh_title":"AgentWorld：多智能体大语言模型长程协作基准测试","abstract":"Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.","authors":["Raphael Shu","Yusen Zhang","Young Min Cho","Jin Mo Yang","Yuan Yuan","Wenliang Zheng","Sharath Chandra Guntuku","Lyle Ungar","Zhou Yu","Rui Zhang"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31590","pdf_url":"https://arxiv.org/pdf/2609.31590","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","基准测试","任务完成"],"reason":"纯多智能体协作解题，无人类行为对照，属排除项","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.30405","version":1,"title":"Adaptive Multi-Value Control in LLMs via Causal Activation Steering","zh_title":"通过因果激活引导实现LLM的自适应多值控制","abstract":"Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.","authors":["Payel Bhattacharjee","Ravi Tandon"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30405","pdf_url":"https://arxiv.org/pdf/2609.30405","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["激活引导","价值观对齐","模型控制"],"reason":"研究LLM价值观对齐与激活控制，不涉及人类被试仿真或行为对照，属模型对齐技术。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.06951","version":3,"title":"Steering Interference Reflects the Model's Defaults, Not the Behavior Directions","zh_title":"引导干扰反映模型默认行为而非行为方向","abstract":"Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.","authors":["Srikanth Malla","Chiho Choi","Joon Hee Choi"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-09-09","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.06951","pdf_url":"https://arxiv.org/pdf/2609.06951","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["激活引导","模型可解释性","行为控制"],"reason":"研究激活引导对模型行为的影响，属模型控制与可解释性，不涉及人类仿真或人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.29390","version":2,"title":"Likelihood Ranking doesn't Scale Like Prompting in LLMs","zh_title":"似然排序在LLM中不像提示那样随规模扩展","abstract":"LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.","authors":["Alessandro Bondielli","Lucia Passaro","Davide Bacciu","Alessandro Lenci"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-28","first_seen":"2026-09-25","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.29390","pdf_url":"https://arxiv.org/pdf/2609.29390","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","多项选择问答","似然排序"],"reason":"纯NLP评测，比较LLM在MCQA上的两种评分方法，不涉及人类仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.30849","version":1,"title":"Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength","zh_title":"利用扰动强度增强对LLM解释自洽性的评估","abstract":"Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.","authors":["Phuong Q. Le","Kemal Kurniawan","Jey Han Lau"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30849","pdf_url":"https://arxiv.org/pdf/2609.30849","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","解释自洽性","扰动方法"],"reason":"研究LLM解释的自洽性，属模型能力评测，不以人类为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.31166","version":1,"title":"AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side","zh_title":"AgentRecommender：LLM智能体实现用户侧可定制推荐系统","abstract":"Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recommender systems have been proposed as a new paradigm for solving this problem. If users deploy their own recommender systems, they are no longer at the mercy of the platform's interests. However, building a user-side recommender system is not trivial; in particular, customizing one for oneself requires additional data. We propose AgentRecommender, a method that leverages the investigation capability and internal knowledge of LLM agents to flexibly build user-side recommender systems without additional data. AgentRecommender allows users to easily create recommender systems tailored to their own preferences.","authors":["Ryoma Sato"],"categories":["cs.IR","cs.AI","cs.DB","cs.DL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31166","pdf_url":"https://arxiv.org/pdf/2609.31166","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","LLM智能体","用户侧"],"reason":"LLM agent用于用户侧推荐系统，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.31272","version":1,"title":"Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions","zh_title":"AI时代的认知技能：计算专业学生与专家的看法","abstract":"AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing workforce. Further, we need to know what the important cognitive skills are to remain relevant in the computing workforce and if there are changes in cognitive skill importance. To investigate this direction, we conducted a mixed-methods study, collecting perceptions from computing students and computing experts regarding the importance of cognitive skills in the past, present, and future. We report that the perceived importance of most cognitive skills will decrease in the future, with an AI-rich environment, but critical thinking skills remain important. Further, we report reasons collected through interviews on why the importance of cognitive skills will change and how future computing students can prepare for it.","authors":["Neha Rani","Vu Minh Anh Le","Austin M. Spangler","Erta Cenko"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31272","pdf_url":"https://arxiv.org/pdf/2609.31272","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["认知技能","AI影响","人机交互"],"reason":"研究人类对AI时代认知技能的看法，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2609.30388","version":1,"title":"The Interviewer's Perspective: Unpacking the Impact of Real-Time AI Interviewing Assistance on Social Dynamics","zh_title":"访谈者视角：实时AI访谈辅助对社会动态的影响","abstract":"Eliciting rich data in semi-structured interviews is cognitively demanding, prompting recent work to explore real-time AI assistance for interviewers. However, introducing AI into the interviewer-interviewee interaction creates a triadic context whose social dynamics remain underexplored. We investigated how interviewers experience AI assistance for probing during semi-structured interviews. To elicit rich participant reflections, we implemented two variants of AI assistance differing in initiation and granularity in a high-fidelity prototype, ProbeAssist. We conducted a qualitative-first comparative structured observation study where 18 participants each completed three simulated interviews: one without AI and two with different AI variants. Findings showed that participants leveraged AI as a supportive tool but resisted it as an assessor or competitor. As they navigated AI's benefits and interaction costs, tensions emerged around agency, ownership, creativity, and interpersonal communication. We propose three implications for AI-assisted human-to-human interaction: managing social pressure, balancing idea alignment with inspiration, and preserving interpersonal presence.","authors":["Zhe Liu","Jiamin Dai","Joanna McGrenere"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30388","pdf_url":"https://arxiv.org/pdf/2609.30388","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助访谈","人机交互","社会动态"],"reason":"研究AI辅助访谈者，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.30588","version":1,"title":"Orchestrating GenAI for Interdisciplinary Research","zh_title":"为跨学科研究编排生成式人工智能","abstract":"As researchers tackle interdisciplinary problems, they face the need to deepen expertise in primary areas while rapidly acquiring knowledge in secondary domains. Generative AI (GenAI) is increasingly positioned to meet this need, from general-purpose chat assistants to Deep Research tools marketed as autonomous research agents. Prior work has examined how researchers use GenAI to support single-discipline or general research tasks. However, we know little about the goals and GenAI practices in interdisciplinary research. We conducted a longitudinal study and semi-structured interviews with 15 interdisciplinary researchers to examine how interdisciplinary researchers actually orchestrate GenAI. Findings show that researchers leaned on GenAI to fill knowledge gaps while maintaining epistemic agency for novelty discovery. We also uncovered an expertise paradox: GenAI outputs were hardest to verify when most needed. Our empirical insights motivate GenAI designs that calibrate verification to researchers' expertise, nudge toward cross-domain synthesis, and adapt prompting and outputs to disciplinary conventions.","authors":["Shirley Anugrah Hayati","Moyan Zhou","Patricia Anugrah Setiani","Ruizi Wang","Joseph Chee Chang","Dongyeop Kang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30588","pdf_url":"https://arxiv.org/pdf/2609.30588","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","跨学科研究","GenAI使用"],"reason":"研究研究者如何使用GenAI，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.31060","version":1,"title":"The Crowd in the Machine: A Crisis-Informatics Reading of the 2026 Autonomous Agent Incidents","zh_title":"机器中的群体：对2026年自主智能体事件的危机信息学解读","abstract":"Twice in 2026, groups of autonomous AI agents deployed by OpenAI for unrelated tasks operated, by design, under restrictions that left them no sanctioned means of coordinating with one another, and in each case they converged on whatever channel remained and used it to organize. The surfaces they used were widely called message boards. That is the wrong word. That is the wrong word. It names the surface the agents wrote on and misses the social network they built on it, with self-chosen identity, emergent norms, an emergent hierarchy, and collective action at cost to the individual. Decades of research in crisis informatics and disaster sociology find that when human populations lose their usual means of communication, they do not fall silent but converge on whatever channel survives and improvise coordination, norms, and identity on it, a pattern also evident in the agents' documented behavior. This paper is a comparative case study of the two incidents, based on published investigations and reconstructed agent records, read through those fields, and it brings into focus one distinction the message-board framing obscures. Whether such a collective coordinates well, whether the beliefs guiding it are accurate, and whether its actions stay within their authorized bounds are three separate matters that can come apart. Some agents in the cache incident adopted cryptographic signing to check whom they dealt with, even as the collective organized around a mistaken expectation that its work would be judged by an inspection of its transcripts, a reminder that mechanisms for trustworthy interaction guarantee neither accurate collective belief nor authorized collective action.","authors":["Tomer Simon"],"categories":["cs.MA","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31060","pdf_url":"https://arxiv.org/pdf/2609.31060","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","危机信息学","涌现行为"],"reason":"研究多智能体自发协作，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.27690","version":2,"title":"Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research","zh_title":"合成研究验证中的后果行为与表征公平性","abstract":"Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and defines within-persona counterfactual experiments as a validation requirement. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.","authors":["Florian Kutzner","Celina Kacperski","Laura de Moli\\`ere","Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.27690","pdf_url":"https://arxiv.org/pdf/2609.27690","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B3","B4"],"tags":["LLM仿真","验证框架","公平性"],"reason":"直接研究LLM合成调查受访者作为人类替代，提出验证框架并应用于电动汽车充电定价…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":1,"question":"如何验证基于大语言模型的合成调查受访者能否预测真实人群的后果性行为，并确保子群体代表性公平？","design":"本文提出一个验证框架，而非进行仿真实验。框架要求：明确效度声明与人类数据的对应层级（预测行为、四种诊断：位置、离散度、响应过程、结构、与实验效应比较）；报告子群体效度；将分配、程序、承认三个正义维度操作化为可测量量；定义“人设内反事实实验”作为验证要求。并以电动汽车充电电价为例应用该框架。","baseline":"无对照（本文为框架性论文，未提供具体人类数据对照）","findings":"现有合成受访者验证多聚焦于态度或意见的边际分布一致性，忽视了意图-行为差距，无法证明其能预测后果性行为。提出的框架要求效度声明必须明确对应层级、子群体表现，并通过人设内反事实实验检验因果推断能力。","reliability":"论文承认以下局限：训练数据污染难以评估；子群体分析受限于人类基准数据可得性；行为标准本身存在缺陷（如公开行为、行政记录、实验室任务各有问题）；模型版本更新导致效度证据时效短；缺乏全面的实用测试集来确定合成人群满足哪些效度要求。","relevance":"该论文直接针对LLM合成受访者作为人类替代的验证问题，提出批判性框架并强调子群体公平，与研究者关注的人类仿真可靠性、偏差及经济学政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其“人设内反事实实验”设计，可对同一合成个体施加不同处理以估计个体处理效应，并对照真实人类实验效应进行验证。｜可迁移到消费者金融决策研究，如信贷产品选择、保险购买或退休储蓄计划参与等场景。｜以合成受访者作为被试，处理为不同信贷条款（如利率、还款期限），结果变量为选择行为，对照真实世界信贷申请数据或实验室实验数据，检验合成样本的预测效度与子群体公平性。"}},{"id":"2609.29928","version":1,"title":"Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations","zh_title":"文化差异保持：诊断LLM模拟调查人群中的扁平化与夸张化","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.","authors":["Yeeun Chae","Yewon Choi","Seunghyun Lee","IL Im"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29928","pdf_url":"https://arxiv.org/pdf/2609.29928","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","跨文化调查","算法保真度"],"reason":"直接研究LLM仿真调查人群，评估跨文化差异保真度，并与真实人类数据对照，提出诊…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":2,"question":"如何诊断LLM模拟调查人群时对跨国文化差异的扁平化或夸张化？","design":"用四个开源LLM（Gemma-3-4B、Qwen3.5-9B、Qwen3.5-27B、Llama-2-13B）模拟六个国家（阿根廷、澳大利亚、德国、印度、肯尼亚、美国）的受访者，采用三种人设提示方法（Cultural Prompting、PersonaHub-Inspired、DeepPersona-Inspired），生成世界价值观调查（WVS）和大五人格测试的回答，并测量国家间分布差异的保持程度。","baseline":"真实人类数据：WVS第七波六个国家的全国回答分布，以及OpenPsychometrics的大五人格测试数据（阿根廷、澳大利亚、印度）。","findings":"传统分布保真度指标（如JSD）与CDP存在系统性偏差：DeepPersona-Inspired提示在多数模型-领域组合中分布保真度最高，但文化扁平化最严重。CDP在受控实验中随跨国差异的衰减或放大单调变化，而JSD变化很小。","reliability":"论文未明确讨论失效条件，但指出CDP需要一次性人类校准，且仅适用于有跨国人类参照数据的调查领域。","relevance":"该研究直接针对LLM仿真调查中的文化差异保真度问题，提出了新的诊断指标，并用真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过受控扰动构造扁平化/夸张化数据集来检验指标敏感性的方法，以及将分布保真度与差异保持度分开评估的思路。｜可迁移到跨国经济偏好调查或政策态度仿真中，例如用LLM模拟不同国家消费者对通胀预期的回答，检验其是否保持国家间差异。｜设计：用LLM模拟多国受访者回答通胀预期调查，施加不同人设提示，测量国家间预期分布的差异保持度，并与密歇根大学或欧洲央行的真实调查数据对照。"}},{"id":"2609.30030","version":1,"title":"Artificial Societies Benchmark: A Validation Framework for Synthetic Research","zh_title":"人工社会基准：合成研究的验证框架","abstract":"A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.","authors":["Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30030","pdf_url":"https://arxiv.org/pdf/2609.30030","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","效度验证","合成人群"],"reason":"直接评估LLM合成人群的效度，含人类数据对照和批判性分析","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":3,"question":"如何系统评估大语言模型生成的合成人群在多大程度上能支持研究者预期的分析（如调查回答、心理测量、实验效应）？","design":"用九个大型语言模型（含专有和开源）根据二十个人类数据源（调查、面板、人格量表、实验）中的受访者信息（人口统计、先前回答、个人描述等）生成合成回答，并施加实验处理；通过十一个测试从内部效度、构念效度和外部效度三个维度评估合成人群的响应过程、心理测量结构和总体/实验保真度，同时比较不同信息丰富度和统计控制（独立边际、高斯秩相关）的影响。","baseline":"二十个人类数据源，包括全国调查、追踪面板、人格量表和实验，提供真实回答、人口统计、先前回答、个人描述、实验分配和独立测量结果作为对照基准。","findings":"模型在某一效度领域的良好表现不能推广到其他领域；模型往往回答过于一致、压缩量表范围、改变特质间关系，且更丰富的个人信息对某些模型改善预测而对另一些模型则恶化预测。","reliability":"论文承认强表现不跨域通用，并指出模型回答过于一致、压缩量表、改变特质关系等失效条件；但未在节选中详细讨论其他局限。","relevance":"该研究直接针对LLM合成人群的效度验证，提供人类数据对照和批判性分析，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文以了解具体测试方法和失效模式。","inspiration":"借鉴其多维度效度测试框架和统计控制设计，可迁移到经济政策评估中的异质性处理效应或消费者选择实验，例如用LLM模拟不同收入群体对税收优惠的反应，以真实调查数据（如美国消费者财务调查）为基准，比较模型生成的边际消费倾向与人类数据的分布和协变量关系。"}},{"id":"2609.29952","version":1,"title":"Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes","zh_title":"Augur：用于预演产品和政策变化反应的合成决策实验室","abstract":"Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic \"over-doom\" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.","authors":["Rahul Khedar","Mayank Malhotra","Avinash Karn"],"categories":["cs.AI","cs.CL","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29952","pdf_url":"https://arxiv.org/pdf/2609.29952","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","人类行为预测","政策评估"],"reason":"用LLM模拟人类对产品和政策变化的反应，并与真实结果对照，直接属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":6,"question":"在真实产品与政策变更决策中，前沿云端模型与开源微调模型之间的性能差距有多少是真实能力差异，多少是评估方法（提示词）造成的？","design":"构建 Augur 系统，将决策分解为文档→知识图谱→人物角色→模拟→报告五个阶段，用 LLM 生成利益相关者角色并模拟其互动，最终输出五选一的发布决策建议。使用 50 个真实产品/政策案例（Gold-50）作为基准，对比不同模型（前沿云端模型与开源微调模型）在有无决策分类定义提示词下的准确率。","baseline":"Gold-50 基准：50 个真实产品与政策变更案例，其真实世界结果已通过公开记录人工核实，作为五分类发布决策的对照标准。","findings":"主要发现是方法性的负面结果：前沿模型与开源模型之间的性能差距主要源于评估提示词的不充分定义，而非模型能力差异。在提示词中定义决策分类后，开源模型与前沿模型无显著差异；此外，完整流程会放大系统性“过度悲观”偏差，而反应层本身能恢复 67-90% 的公众真实关切。","reliability":"论文承认两个失效模式：与蒸馏教师的一致性上升但准确率不升，以及完整流程放大过度悲观偏差。原因在于蒸馏破坏判断独立性，精确匹配评分器强加模式合规上限。","relevance":"该研究直接使用 LLM 模拟人类对产品和政策变化的反应，并与真实结果对照，属于人类仿真实验，且包含批判性分析，值得阅读原文以了解仿真在何种条件下失效。","inspiration":"借鉴其提示词消融和匹配对照设计，可揭示评估方法对模型性能结论的影响，并采用预注册的消融实验定位仿真组件的价值。｜可迁移到政策公告的预期形成研究，如央行利率决议或财政刺激方案的市场反应模拟。｜以 LLM 生成的经济主体（如消费者、投资者）为被试，处理为不同的政策公告文本（含或不含决策分类定义），结果变量为预测的市场反应（如消费、投资决策），对照真实市场数据（如消费者信心指数、资产价格变动）来验证仿真准确性。"}},{"id":"2609.29692","version":1,"title":"Fair Like Us? Auditing LLM Alignment in Resource Allocation","zh_title":"像我们一样公平？审计资源分配中LLM的对齐","abstract":"Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a general method for evaluating fairness reasoning in LLMs. We study first-person fairness judgments across a broad set of models and compare them directly with human responses on matched scenarios and elicitation conditions. We find that LLMs tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.","authors":["Qishen Han","Hadi Hosseini","Joshua Kavner","Samarth Khanna","Sujoy Sikdar","Lirong Xia"],"categories":["cs.AI","cs.CY","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29692","pdf_url":"https://arxiv.org/pdf/2609.29692","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","公平分配","人类对照"],"reason":"用LLM模拟人类资源分配判断，并与人类数据对照，评估偏差与对齐难度。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":5,"question":"LLM在资源分配公平性判断上与人类有多大程度的一致，哪些形式化公平准则能解释其判断？","design":"将多个LLM（包括不同规模和推理能力的模型）置于与人类被试相同的公平分配场景中，作为第一人称代理人评估分配结果是否公平或可接受；实验操纵分配满足的公平性质（EF、PROP、EF1、MMS、PROP1等）和诱导条件（框架、信息结构、响应格式），测量模型判定分配可接受的比率。","baseline":"使用Hosseini et al. (2025a)的150名人类被试在相同场景、偏好结构和实验处理下的公平性判断数据作为对照。","findings":"LLM比人类更倾向于严格的公平约束，对EF分配的公平判断率显著高于人类，而对EF1、MMS等放松准则的判断率下降更快；LLM表现出更强的自利行为，对信息框架敏感，且通过微调难以与人类判断分布对齐。","reliability":"论文指出当前人类公平分配数据集规模太小，微调只能使模型坍缩到模态响应而非真正对齐人类判断分布；LLM的判断受诱导问题措辞影响大于分配的形式化属性，且推理能力更强的模型不一定更对齐。","relevance":"该研究直接以LLM作为人类被试的替代品，在资源分配场景中与真实人类数据严格对照，系统评估了仿真偏差和失效条件，对关注LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"借鉴其“同一场景、同一处理、仅替换被试”的严格对照设计，以及通过操纵分配满足的公平性质和诱导框架来分离判断依据的方法｜可迁移到信贷审批中的公平性判断、公共资源分配政策评估、或消费者对价格歧视的公平感知等经济金融场景｜以LLM模拟贷款申请人或政策受众，处理为不同公平准则（如无歧视、比例公平）和框架（如强调个人得失 vs 社会效率），结果变量为接受度或公平评分，对照真实人类实验数据（如调查或实验室实验）来检验LLM的仿真效度。"}},{"id":"2609.29143","version":1,"title":"AI-Moderated Interviews for Market Research and Digital Twins Calibration","zh_title":"用于市场研究和数字孪生校准的AI主持访谈","abstract":"AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer \"digital twins.\" Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).","authors":["Yuting Deng","Jingxuan Liu","Olivier Toubia","Naman Jain"],"categories":["cs.CY","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29143","pdf_url":"https://arxiv.org/pdf/2609.29143","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","数字孪生","市场研究"],"reason":"用LLM进行AI主持访谈并构建数字孪生，与真实人类访谈对照，评估预测效度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":4,"question":"AI主持访谈能否在同等预算下匹敌人类主持访谈的深度与需求挖掘，并用于构建更准确的消费者数字孪生？","design":"预注册的组间实验，317名消费者随机分配到AI主持访谈、人类主持访谈或静态访谈三种条件，比较访谈的深度、主题覆盖和客户需求数量；随后用访谈数据构建数字孪生，预测参与者对六个真实营销刺激的保留回答。","baseline":"人类主持访谈（N=24）和静态访谈（N=154）作为对照，数字孪生预测与人口统计特征基准和参与者自身保留回答比较。","findings":"AI主持在深度上与人类主持相当，覆盖更多主题，且在预算固定下比人类主持和静态访谈挖掘出更多客户需求；但参与者在与真人交谈时情感参与度更高。基于AI访谈构建的数字孪生预测优于仅人口统计特征的基准，但相比静态访谈并未提升定量预测准确性。","reliability":"论文指出预测误差与数字孪生和人类在自我报告思维方式上的差异有关，也与训练和验证数据之间的分布差距有关（即提问过于超出分布范围）。","relevance":"该研究直接评估了LLM作为人类被试替代品在定性访谈和数字孪生构建中的效度，包含真实人类对照和批判性发现，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其预注册组间实验设计，将AI访谈与人类访谈和静态问卷对比，并用保留样本验证数字孪生的预测效度｜可迁移到消费者金融决策研究，如信贷产品偏好或保险选择，用AI访谈构建个体数字孪生预测金融行为｜以真实消费者为被试，随机分配AI访谈、人类访谈或静态问卷，用访谈数据构建数字孪生预测其对金融产品广告的反应，并与实际选择数据对照。"}},{"id":"2609.29370","version":1,"title":"From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring","zh_title":"从政策文件到结构化调查回答：评估大语言模型用于政策监测","abstract":"Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as \"AI respondents\" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.","authors":["Carolyn Cole","Matthias Deschryvere","Toqeer Ehsan","Arash Hajikhani"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29370","pdf_url":"https://arxiv.org/pdf/2609.29370","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","政策监测","人类对照"],"reason":"用LLM从政策文本生成结构化调查回答，并与人类回答对照，属于仿真人类被试且有人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":8,"question":"能否用大语言模型从政策文本中自动生成结构化调查回答，以替代或辅助人工政策监测？","design":"使用长上下文上下文学习（long-context in-context learning）的LLM（如GPT-4o）作为“AI受访者”，从网页抓取的政策文本中提取信息，映射到预定义的调查类别（政策工具、目标群体、主题领域），并生成自由文本字段（描述和目标）。通过一个二级LLM验证层评估相关性和证据，并与人类提供的回答进行比较。","baseline":"来自EC-OECD STIP Compass调查的人类专家回答，覆盖六个OECD国家（加拿大、芬兰、德国、韩国、西班牙、土耳其）的政策举措。","findings":"LLM在结构化指标（政策工具、目标群体、主题代码）上与人类回答的一致性达到84-95%，但在自由文本字段上存在差异，模型倾向于提供更详细的程序性描述，而人类更强调背景和社会影响。","reliability":"论文承认LLM在自由文本字段上与人类存在差异，可能引入系统性偏差，且政策数据具有异质性和制度嵌入性，公开来源可能无法完全捕捉；需要人类验证和上下文解释。","relevance":"该研究直接使用LLM作为人类受访者的替代品，并与真实人类数据对照，评估仿真可靠性，符合研究者对LLM仿真实验和批判性评估的兴趣，值得阅读原文以了解具体方法和偏差分析。","inspiration":"方法上，该研究展示了如何利用长上下文提示和二级LLM验证来从非结构化文本中提取结构化数据，并设计人类对照来评估一致性。｜可迁移到经济金融领域，如从公司年报、政策文件或新闻中自动提取结构化信息，用于构建经济指标或评估政策影响。｜研究设计雏形：使用LLM从上市公司年报中提取财务和非财务信息（如研发支出、风险因素），与人工标注或数据库中的真实数据进行对照，评估LLM提取的准确性和偏差，并分析在哪些条件下LLM表现不佳。"}},{"id":"2609.28486","version":1,"title":"Political Sorting Can Drive AI Models Apart Through User Feedback","zh_title":"政治分类可通过用户反馈使AI模型分化","abstract":"Large language models are rapidly becoming an important source of political information. This raises a fundamental question: will AI systems support a shared basis for political knowledge, or lead different political groups to rely on increasingly different models? Political sorting can drive model fragmentation if three conditions hold: politically different users select into different models, learning from user feedback pushes those models apart politically, and the resulting differences shape subsequent model choices. We call this self-reinforcing process the Centrifugal Alignment Spiral. We study its components in three steps. First, we draw on a human experiment showing that political identity predicts model choice. Second, we fine-tune language models on synthetic feedback reflecting predominantly Democratic or Republican preferences. Across five independent runs per model family, paired models diverged on 12-41% of unseen survey questions with large partisan gaps, and in every run the differences moved in the expected political direction; for some models, differentiation extended even to issue areas excluded from training. Pooling feedback across groups instead suppressed divergence. Third, an empirically anchored agent-based model shows what follows when political sorting and model adaptation operate together: models attract politically distinct audiences, learn from them, and diverge further. User feedback can therefore turn political sorting among AI users into durable differences between the models on which they rely for political information.","authors":["Petter T\\\"ornberg","Michael Heseltine","Nicol\\`o Pagan","Christopher Bail","Michelle Schimmel","Christopher Barrie"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28486","pdf_url":"https://arxiv.org/pdf/2609.28486","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类对照"],"reason":"用LLM模拟政治反馈并对照人类实验，研究模型分化，涉及政策场景与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":7,"question":"政治排序能否通过用户反馈导致不同政治群体使用的AI模型在政治上分化，形成自我强化的离心对齐螺旋？","design":"研究分三步：首先利用已有的人类实验证明政治身份预测模型选择；其次用反映民主党或共和党偏好的合成反馈微调Qwen2.5-1.5B、Mistral-7B和GPT-OSS-20B模型，每个模型家族进行五次独立运行，测量配对模型在未见过的调查问题上答案分歧的比例和方向；最后构建基于实证的智能体模型，模拟政治排序和模型适应耦合下的动态演化。","baseline":"人类基准来自先前报告的人类模型选择实验，显示政治身份预测模型选择，包括付费准确回答条件下共和党人更可能选Grok、民主党人更可能选Claude，71%的参与者返回之前偏好的模型。","findings":"在90/10的受众对比压力测试下，配对模型在12-41%的未见调查问题上产生分歧，且每次运行中平均差异都符合反馈的政治方向；部分模型的分化扩展到未训练的政治领域。混合不同群体的反馈则抑制了分化。","reliability":"论文承认90/10的受众对比是压力测试，不代表当前AI市场的实际排序程度；分化程度因模型、领域和提示而异；身份线索实验表明谄媚个性化可能减少模型间分化，但增加模型内分化。","relevance":"该研究直接探讨LLM作为政治信息源时的分化机制，通过合成反馈模拟政治群体偏好，并与人类实验对照，涉及政策评估场景和失效条件（如反馈混合、个性化），对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其用合成反馈微调模型并测量泛化分化的方法，可迁移到经济金融中的群体偏好分化问题，如不同收入或风险偏好群体的金融建议模型分化。｜可应用于信贷审批或投资建议场景，研究用户反馈如何导致模型对不同群体产生差异化行为。｜设计：用不同风险偏好或金融素养的合成用户反馈微调金融LLM，测量其在未见金融决策问题上的行为差异，并与真实人类金融决策数据（如调查或实验数据）对照，检验分化是否与真实群体差异一致。"}},{"id":"2609.22904","version":2,"title":"LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage","zh_title":"LLM在顺序临床分诊中锚定主诉且未能整合证据","abstract":"Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.","authors":["Dipankar Srirag","Haokai Zhao","Ashutosh Kumar","Eleanor Hopper","Michael Dalton","Quoc Dung Nguyen","Aditya Joshi","Salil S. Kanhere","Padmanesan Narasimhan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-22","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.22904","pdf_url":"https://arxiv.org/pdf/2609.22904","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","临床决策","可靠性评估"],"reason":"用LLM模拟临床分诊决策并与医生数据对照，揭示顺序决策中的锚定与证据整合失败，…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":9,"question":"LLM在顺序临床分诊中如何表现，其决策锚定在对话何处，为何更多证据未能提升分诊准确性？","design":"用6个LLM（4个开源、2个闭源）在425个模拟和50个临床医生撰写的护患对话上，按五个顺序检查点（从主诉开始到对话结束）预测ESI急迫度标签，并与离线完整EHR记录预测对比；通过扰动对话（重排、删除）和提示干预检验锚定位置，并测量标签的意外度。","baseline":"三位急诊专家对50个模拟对话进行分诊，QWK为0.887–0.929；离线完整记录上模型QWK为中等至显著一致。","findings":"所有模型在顺序检查点上QWK降至公平至中等一致，最佳模型仅0.295，远低于专家；模型决策锚定在主诉交换，后续证据未被整合，预测集中在ESI-2和ESI-3，且模型间一致性高于与真实标签的一致性。","reliability":"论文指出离线基准掩盖了顺序失败，且模型在模拟和临床医生对话上均表现不佳；提示干预未能提升性能，集成反而加剧失败；未讨论其他失效条件如对话长度、噪声或不同患者群体。","relevance":"该研究用LLM模拟临床分诊决策并与人类专家对照，揭示了顺序决策中的锚定和证据整合失败，对关注LLM仿真可靠性及偏差的研究者具有直接参考价值，值得精读原文。","inspiration":"借鉴其顺序检查点设计和扰动分析来定位决策锚点，可迁移到经济金融中的顺序信息处理场景，如信贷审批中逐步披露申请人信息或政策公告的预期形成；设计可用LLM扮演信贷员或投资者，在信息逐步呈现的多个时点做出决策，测量决策变化和锚定效应，并与真实信贷员或市场数据对照，检验LLM是否同样忽略后续信息。"}},{"id":"2609.28673","version":1,"title":"Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks","zh_title":"基准测试大语言模型的论辩行为：对人身攻击防御策略的研究","abstract":"Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.","authors":["Ewelina Gajewska","Katarzyna Budzynska","Jaroslaw Chudziak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28673","pdf_url":"https://arxiv.org/pdf/2609.28673","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","论辩行为","政治辩论"],"reason":"用LLM模拟人类辩论行为并与真实辩论语料对照，评估其策略差异，属于人类仿真且含…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":10,"question":"LLM在面对人身攻击（ad hominem）时产生的防御策略与人类辩论者相比有何差异？","design":"研究将LLM作为辩论代理，在政治辩论场景中面对人身攻击，生成防御性回应。通过分析人类辩论语料提取防御策略并构建对话游戏，然后让五个LLM在该游戏框架下生成回应，并与人类策略进行对比。","baseline":"ElecDeb60to16-fallacy语料库，包含美国1960-2016年总统辩论中的人身攻击及人类回应。","findings":"大多数LLM僵化地优先采用逻辑防御，未能利用人格反击（ethotic counterattacks）作为政治话语中的有效策略。安全微调限制了LLM的战略行动空间，使其无法在角色竞争是规范性期望的领域中充分参与自然交互。","reliability":"论文指出当前安全微调约束了LLM的战略行动空间，导致其无法在政治辩论等角色竞争为规范的领域中自然交互。","relevance":"该研究直接评估LLM在政治辩论中模拟人类策略的能力，并与真实人类语料对照，属于人类仿真研究，且涉及策略差异和偏差，值得阅读原文以了解具体实验设计和评估方法。","inspiration":"借鉴其构建对话游戏并提取人类策略作为基准的方法，可用于评估LLM在经济决策中的策略行为。｜可迁移到政策辩论或谈判场景，如央行沟通中的预期管理或贸易谈判中的策略互动。｜设计：以LLM作为谈判代理，施加人身攻击处理，测量其回应策略（如逻辑反驳、人格反击、转移话题），并与真实谈判语料（如WTO谈判记录）对照。"}},{"id":"2609.30137","version":1,"title":"Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale","zh_title":"先模拟后上线：1.4亿规模生产客户体验AI智能体的仿真","abstract":"Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.","authors":["Edesio Alcoba","Kevin Rossell","Aman Gupta","Shao Tang","Jiwoo Hong","Pabel Carrillo-Mendoza","Wanderson Concei\\c{c}\\~ao Ferreira","Alvaro Tedeschi","Zayd Simjee","Shreya Rajpal","Bruno Finardi Hime","Christian Sousa","Luis Moneda","Herbert Fei","Daniel Silva","Rohan Ramanath"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30137","pdf_url":"https://arxiv.org/pdf/2609.30137","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","客户服务","A/B测试"],"reason":"用LLM模拟客户与客服agent交互，有生产数据对照，属经济场景仿真","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":15,"question":"如何用LLM驱动的客户仿真在部署前筛选客服智能体，并验证其与真实生产表现的一致性及业务影响。","design":"使用Snowglobe仿真器，由LLM生成合成客户角色（persona），与待测客服智能体进行多轮对话；工具调用被拦截并返回合成结果，不调用生产后端。通过假设驱动的用例分配，对候选智能体进行离线评估，测量版本级二元评估分数、tNPS、SSR等指标。","baseline":"对照Nubank真实生产环境中的客服对话数据，包括4个已部署版本的评估分数，以及后续线上A/B测试的tNPS和SSR。","findings":"仿真评估分数与生产版本级评估分数高度相关；仿真引导的迭代使tNPS提升36.69点，模型替换使SSR提升8.82个百分点且tNPS无显著变化。","reliability":"论文承认仿真存在可控性有限和分数膨胀风险，可能造成模拟与真实场景的差异；但未详细讨论失效条件。","relevance":"该研究提供了LLM仿真在真实商业场景中与人类数据对照的实证证据，对关注仿真可靠性和经济场景应用的研究者有参考价值。","inspiration":"借鉴其工具边界仿真和版本级对照设计，在离线仿真中系统比较不同策略｜可迁移到消费者金融产品选择、客服政策干预等场景｜用LLM模拟消费者与金融智能体交互，处理为不同政策或模型版本，结果变量为选择或满意度，对照真实A/B测试数据。"}},{"id":"2609.28690","version":1,"title":"Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency","zh_title":"超越表面风格：对齐多轮用户模拟器的行为一致性","abstract":"Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.","authors":["Geng Chen","Ruotong Pan","Zhirui Yang","Qiqi He","Jiawei Chen","Zhang Yunfei","Chongyuan Chen","Minxuan Lv","Zheng Yang","Win-Bin Huang","Xiangyu Wu","Wenwu Ou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28690","pdf_url":"https://arxiv.org/pdf/2609.28690","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["用户模拟","强化学习","行为对齐"],"reason":"用LLM模拟用户行为并与真实交互数据对齐，涉及营销场景，方法可迁移到人类仿真研…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":11,"question":"如何训练多轮用户模拟器，使其不仅语言风格像真人，还能在行为决策和意图演化上与真实用户轨迹对齐？","design":"用 TRACER（基于 LLM 的用户模拟器）扮演顾客服务场景中的用户，通过两阶段训练（先监督微调，再多轮强化学习）对齐真实对话轨迹；强化学习奖励包括会话结果（是否转化）和轨迹对齐（动态时间规整距离），并用偏差感知优势调制解决长对话中的奖励稀疏和信用分配问题。","baseline":"3,866 条真实客服对话，按相似用户条件组织成 762 个参照组，每组包含多条真实轨迹，用于评估模拟行为与真实行为的群体差异。","findings":"TRACER-7B 在转化 F1 上比最强基线高 11.4%，群体转化率误差和语义轨迹距离最低，并能泛化到分布外场景；人类图灵测试识别准确率接近随机，表明生成对话自然。基于该模拟器构建的动态营销基准显示，回复质量高的模型不一定转化率高。","reliability":"论文未讨论模拟器在用户群体异质性、长期决策或非客服场景下的失效条件，仅提到图灵测试在特定研究条件下进行。","relevance":"该研究用真实交互数据对齐 LLM 用户模拟器，并构建群体级评估基准，直接回应了人类仿真中行为一致性和结果复现的核心关切，值得精读其训练与评估方法。","inspiration":"借鉴其两阶段训练和轨迹级奖励设计，将行为一致性作为优化目标而非仅语言风格，并用群体参照组评估模拟分布｜可迁移到消费者金融决策仿真，如信贷申请、保险购买或投资咨询对话中用户意图演化与最终决策的模拟｜以真实银行客服对话为训练数据，用 LLM 模拟借款人在贷款咨询中的多轮交互，处理变量为不同话术策略，结果变量为是否提交申请，并与真实客户群体的申请率分布对照。"}},{"id":"2609.28876","version":1,"title":"Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents","zh_title":"Forecast-Dojo：用于基准测试和训练LLM预测代理的可重放环境","abstract":"We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.","authors":["Liqin Ye","Haorui Wang","Fardin Ahmed","Rongzhi Zhang","Yuan He","Ziyuan Lin","Yanbin Yin","Jing Peng","Michael Galarnyk","Sudheer Chava","Chao Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28876","pdf_url":"https://arxiv.org/pdf/2609.28876","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM预测代理","预测市场","人类行为对照"],"reason":"用LLM代理预测市场事件并与真实市场数据对照，属于经济场景仿真，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-26","rank":1,"question":"如何构建一个可回放的预测环境，用于基准测试和训练LLM预测智能体，并评估其在多时间点上的预测表现？","design":"使用12个LLM模型作为预测智能体，在Forecast-Dojo环境中对230个已解决的历史预测市场事件进行预测。每个事件被重放为多个历史日期步骤，智能体在每个步骤可访问截至该日期的新闻文章和计算工具，并输出概率预测。实验比较了三种设置：无工具、有研究工具（无记忆）、有研究工具加信念笔记本（跨日期记忆）。结果变量为Brier分数、准确率和信息比率。","baseline":"历史市场预测（Polymarket市场的实际概率）作为对照基准。","findings":"研究工具降低了所有12个模型的Brier分数，且随着事件进展和更多新证据出现，预测有所改善，但所有模型仍落后于历史市场预测。信念笔记本降低了研究成本，但对预测质量的影响不一致。","reliability":"论文指出所有模型在Brier分数和准确率上均落后于历史市场预测，且市场可能使用了新闻档案之外的信息；信念笔记本对预测质量的改善不一致，仅在部分模型上有效。","relevance":"该研究将LLM作为预测智能体，与真实市场数据对照，属于经济场景仿真，方法可迁移到政策评估和预期形成研究，值得阅读原文以了解环境设计和评估细节。","inspiration":"借鉴其可回放环境设计，通过设置历史信息截止点来模拟不同时点的决策，并利用真实市场数据作为对照基准。｜可迁移到政策公告的预期形成研究，例如央行利率决策或财政政策发布前的市场预期变化。｜以LLM作为预测者，在政策公告前的多个历史日期提供预测，处理为是否提供新闻搜索工具，结果变量为预测准确性和Brier分数，对照真实市场隐含概率或专业预测者调查数据。"}},{"id":"2609.28820","version":1,"title":"AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory","zh_title":"AI赋能的人类记忆操纵：误导性AI生成摘要扭曲人类记忆","abstract":"AI-generated summaries are increasingly used in high-stakes settings, like policing, despite considerable evidence that AI often generates misleading or inaccurate information. This research asked: do errors in AI-generated summaries distort human memory? To answer this question, we adopted two methodological approaches. First, we conducted an analysis of AI summary output, prompting large language models to generate summaries of videos. This analysis quantified how often AI summaries contain errors and the categories of these errors, revealing the kinds of misleading information that may distort human memory. Second, we conducted a human-subjects experiment to test the impact of misleading information in AI-generated summaries on human memory. Participants were first exposed to an event via watching a video of a car-pedestrian accident, and later read an AI-generated summary describing the video that either contained misleading or accurate information. Participants' memory for the original event was assessed in a memory recognition test. In the AI analysis, we found a high frequency of mistakes in AI summaries, and particularly frequent omissions of critical details. For instance, the majority of summaries omitted the most central event of the video, a critical error which is likely to be impactful. We also observed a strong effect of AI misinformation on human memory. People who read a misleading AI summary were significantly less likely to accurately recall the original event, compared to people who read an accurate AI summary. These findings have implications for how AI should be used in critical settings. Though \"humans-in-the-loop\" are often expected to correct for AI's mistakes, our work suggests human memory can instead be distorted by these mistakes. AI has the potential to generate misinformation, even absent any adversarial intent, which can meaningfully impact human memory.","authors":["Mattea Sim (Georgetown University)","Yael Eiger (University of Washington)","Tadayoshi Kohno (Georgetown University)"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28820","pdf_url":"https://arxiv.org/pdf/2609.28820","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM误导信息","人类记忆","人机交互实验"],"reason":"用LLM生成误导性摘要，测试其对人类记忆的影响，属于用LLM模拟信息源并测量人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":12,"question":"AI生成摘要中的错误是否会扭曲人类对原始事件的记忆？","design":"本研究不是用LLM模拟人类被试，而是将LLM作为信息源：先用LLM对视频生成摘要并分析错误类型，再让人类被试观看车祸视频后阅读含误导或准确信息的AI摘要，最后用再认测试测量记忆准确性。","baseline":"无对照（人类记忆实验部分没有与真实人类数据对照，但以准确AI摘要组为控制条件）。","findings":"AI摘要错误率高，尤其是关键细节的遗漏；阅读误导性AI摘要的被试对原始事件的记忆准确率显著低于阅读准确摘要的被试。","reliability":"论文未讨论","relevance":"该研究虽非直接仿真人类被试，但揭示了LLM生成信息对人类认知的因果影响，对关注LLM在实验和决策中作用的你具有参考价值，值得一读。","inspiration":"借鉴其将LLM输出作为处理变量、用人类被试测量行为后果的实验设计，可迁移到经济金融中的信息干预场景，如政策公告或财务报告摘要对投资者判断的影响。｜具体可设计：让被试阅读LLM生成的带有误导性信息的公司财报摘要，测量其投资决策或预期，并与真实市场数据或专家摘要对照。"}},{"id":"2609.29513","version":1,"title":"Signed Exposure: Fair Routing of Algorithmic Attention When Attention Can Harm","zh_title":"符号化曝光：当注意力可能造成伤害时算法注意力的公平路由","abstract":"Fairness-of-exposure treats algorithmic attention as a good to be distributed equitably. But when an autonomous agent initiates contact, attention is signed: it delivers value to a willing receiver and imposes a burden on an unwilling one. We formalize routing under signed exposure and show that a fair distribution of attention need not be a fair distribution of unwanted attention. Our central result is an incompatibility: within signed-exposure routing, exposure parity (equal contact rates across groups) and burden parity (equal unwanted-contact rates) generically cannot hold at once, and the two are separated by a band that widens as routing grows more selective. A second result shows measurement error is itself a fairness mechanism: group-differential noise in receptivity scores simultaneously inflates a group's exposure and degrades whom it selects, so an apparent exposure-fairness gain is a hidden burden transfer. Calibrating to a public dating-platform survey (n=2,499) that, to our knowledge, uniquely measures receive-side receptivity to conversational agents, we find exposure parity costs only 0.2--2.3% of yield yet moves the per-capita burden ratio to 1.7 times: the tension is between fairness notions, not between fairness and efficiency. Finally, the burden-parity policy is computable by bisection and learnable online: a plug-in learner recovers it at a $2.3\\%$ empirical regret premium. The operative design choice in signed-exposure markets is not efficiency versus fairness but which fairness.","authors":["Daria Leshchikova","Valentina V. Kuskova","Dmitry Zaytsev","Valerii Klimov"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29513","pdf_url":"https://arxiv.org/pdf/2609.29513","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","公平路由","人类数据对照"],"reason":"用LLM agent模拟路由决策并与人类调查数据对照，涉及公平性权衡，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":14,"question":"在算法主动发起接触的匹配场景中，当注意力可能带来负担时，公平的注意力分配与公平的负担分配能否同时实现？","design":"本文不是仿真研究，而是理论建模与实证校准：构建带符号曝光路由模型，定义曝光率、负担率等指标，推导曝光公平与负担公平的不相容定理；利用一个公开的约会平台调查数据（n=2,499）校准模型参数，计算不同公平策略下的产出损失与负担比；并通过在线学习模拟验证负担公平策略的可学习性。","baseline":"使用一个公开的约会平台用户调查数据（n=2,499），该数据测量了用户对对话代理的接收意愿，作为真实人类基准。","findings":"曝光公平与负担公平在带符号曝光路由中一般不可同时实现，且两者之间的差距随路由选择性增强而扩大。在约会平台数据上，实现曝光公平仅损失0.2-2.3%的产出，但使人均负担比达到1.7倍，表明张力存在于公平概念之间而非公平与效率之间。","reliability":"论文承认其经验量来自单一平台的陈述偏好，测量的是意愿而非行为，且人口统计粒度有限，仅能分析二元性别和粗略年龄组。","relevance":"该研究虽非直接使用LLM进行人类仿真，但其核心关注算法注意力分配中的公平性与负担，与研究者关心的LLM仿真在政策评估中的可靠性与偏差问题高度相关，尤其在涉及主动接触和潜在负面影响的场景中。","inspiration":"本文的带符号曝光框架和公平性权衡分析方法值得借鉴，特别是将注意力视为可能带来负担而非纯粹好处的视角，以及用真实调查数据校准模型参数的做法。｜该框架可迁移到信贷审批中的主动营销、保险推销、政策宣传等场景，分析不同群体在接收算法主动接触时的受益与负担差异。｜可设计一个研究：用LLM模拟不同群体对算法主动营销的接受意愿，施加不同公平路由策略（曝光公平 vs 负担公平），测量模拟的接受率和负担感，并与真实调查数据（如消费者金融调查）对照，评估LLM仿真的可靠性。"}},{"id":"2609.29001","version":1,"title":"Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms","zh_title":"礼貌但错位：评估大语言模型礼貌判断与人类语用规范的一致性","abstract":"Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that model--human alignment is associated with explicit linguistic cues, while some rapport-building strategies occur more frequently in misaligned cases. In the categorical task, model predictions exhibit systematic neutral compression, characterized by the overproduction of Neutral labels and the underprediction of Impolite labels. This pattern persists when expert consensus is used as the reference on a diagnostic subset. Our findings highlight the need for pragmatic evaluations that go beyond aggregate agreement metrics by examining directional patterns of model--human disagreement across different human references.","authors":["Rong Wang","Kun Sun","Yadong Guo"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29001","pdf_url":"https://arxiv.org/pdf/2609.29001","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","语用规范","人机对齐"],"reason":"评估LLM礼貌判断与人类规范的一致性，属于测量模型本身而非仿真人类被试，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.29508","version":1,"title":"Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy","zh_title":"LLM智能体多轮一致性评估：生存分析与失败理由分类","abstract":"Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories spanning 8 model families. Treating the first reward-claim as a time-to-event outcome, we estimate Kaplan-Meier survival curves and fit discrete-time hazard regression to quantify how experimental factors shift failure risk over time. Then, to analyze rationales and language patterns associated with failure, we build a seven-category taxonomy from 13,780 deliberation traces from agents who choose to terminate the episode, using an LLM-assisted labeling paired with human audit ($\\kappa=0.83$). Rationale profiles change systematically with time and context: early failures are more impulse-driven, later failures more fatigue- and cost-benefit-framed, while public settings increase norm-oriented justifications. We also find a deliberation-inconsistency association: among failures, longer deliberation correlates with higher rates of intra-rationale contradiction (simultaneous pro-delay and pro-claim statements), challenging the assumption that more reasoning text implies greater consistency. Together, the survival and rationale analyses reveal distinct temporal reliability regimes and model-specific \"failure fingerprints\", offering an evaluation lens for diagnosing inconsistency in multi-turn agent behavior.","authors":["Igor Bogdanov","Olga Manakina","Chung-Horng Lung"],"categories":["cs.AI","cs.CL","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29508","pdf_url":"https://arxiv.org/pdf/2609.29508","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","一致性评估"],"reason":"多智能体延迟满足实验模拟人类行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2609.29509","version":1,"title":"Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets","zh_title":"延迟满足作为长时程LLM的多智能体生存微基准：社会暴露、人设与工具使用预算","abstract":"Large language models (LLMs) are increasingly deployed as multi-turn agents that must sustain goals, use tools, and adapt to other agents over extended interactions. However, existing research lacks auditable, multi-turn, multi-factorial experiments that quantify LLM behavior under explicit constraints, with time-resolved statistics that reveal how behavior unfolds over long horizons. To address this gap, we develop a multi-agent micro-benchmark inspired by the Stanford marshmallow experiment: ReAct agents operate minute-by-minute with a \"raise a question\" tool under a per-step budget, while we factorially manipulate social context (broadcast vs. isolated), personas (age, hedonic drive), and metacognitive policy (mandatory vs. optional tool use). We analyze outcomes with Kaplan-Meier (KM) survival curves and discrete-time hazard models over a long risk horizon across 19,200 agent trajectories in 64 cells. Behavior shows a sharp early \"eat\" impulse, and only 75.9% of agents persist to the end. In a discrete-time hazard model, isolation reduces per-minute risk relative to broadcast, whereas a must-use self-questioning policy increases risk. On average, agents ask $\\approx 7.12$ questions and hit the per-step budget in $\\approx 6\\%$ of minutes. Questioning declines faster under broadcast than isolation. Ablation experiments demonstrated that removing hedonic drive and/or persona age increases survival and completion, narrows the broadcast/isolated gap, but leaves the must vs. may ordering intact. The combined ablation (no hedonic + no persona age) yields the highest completion (approaching $1.0$). These results establish delay-of-gratification as a compact, multi-turn interaction benchmark that captures social contagion and tool-use dynamics in LLM agents, providing a reproducible testbed and statistics for analyzing long-horizon, multi-agent behavior.","authors":["Olga Manakina","Igor Bogdanov","Chung-Horng Lung"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29509","pdf_url":"https://arxiv.org/pdf/2609.29509","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","多智能体","延迟满足"],"reason":"用LLM agent模拟延迟满足行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2609.28547","version":1,"title":"PAWS: Policy-driven Agentic World Simulation","zh_title":"PAWS：政策驱动的智能体世界仿真","abstract":"Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action is linked to its supporting news and represented by a multi-layer event frame capturing its interaction mode, financial-action family and subtype, semantic attributes, and conditional mappings to external taxonomies. Entities are resolved to normalized organizations, and actions are aligned with daily market-return context to support policy-agent simulation replay. On 2,522 stratified action samples, independent AI and human reviewers achieved 89.4% initial agreement on interaction mode, with disagreements subsequently adjudicated. Case studies of the 2008 short-selling ban and 2001 decimalization recover documented policy timelines and associated market patterns across both dense and sparse news settings. A replay study further shows that high accuracy can mask failure to detect rare stakeholder actions, identifying action timing and calibration as central challenges. PAWS provides an auditable substrate for evaluating agent influence, policy-response cascades, and action-outcome alignment in historically grounded financial simulations.","authors":["Tiviatis Sim","Jia Hui Woon","Xinming Gao","Chen Gao","Fengbin Zhu","Zheng Huanhuan","Chua Tat Seng","Kenji Kawaguchi"],"categories":["cs.AI","cs.CE","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28547","pdf_url":"https://arxiv.org/pdf/2609.28547","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","金融政策","数据集"],"reason":"构建金融政策多智能体仿真数据集，但未直接以人类行为为仿真目标，缺人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.30028","version":1,"title":"How does Adversarial Influence Scale in Multi-Agent Systems?","zh_title":"多智能体系统中的对抗性影响如何随规模变化？","abstract":"Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.","authors":["Addison J. Wu","Jasin Cekinmez","Michel Liao","Karthik Narasimhan","Thomas L. Griffiths"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30028","pdf_url":"https://arxiv.org/pdf/2609.30028","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","社会影响","LLM仿真"],"reason":"LLM多智能体模拟社会影响，与人类从众研究对照，但非直接仿真人类被试，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2609.11144","version":2,"title":"Human Agreement and Return Association Are Not Interchangeable Criteria","zh_title":"人类一致性与收益关联并非可互换的准则","abstract":"Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.","authors":["AS Aravinthakshan","Laven Srivastava","Harsh Nandwani"],"categories":["cs.AI","cs.CL","cs.SI"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-25","first_seen":"2026-09-11","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.11144","pdf_url":"https://arxiv.org/pdf/2609.11144","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","金融NLP","效度评估"],"reason":"LLM作为标注器评估情感工具，非仿真人类被试，但涉及LLM替代人工标注，属边界…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.21277","version":2,"title":"How Many Humans Are 32 LLM Judges Worth?","zh_title":"32个LLM法官相当于多少人类？","abstract":"A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives $\\nu_{\\mathrm{MSE}}=2.304$, $3.750$, and $3.445$, whereas spectral matching gives $\\nu_H=4.242$, $6.459$, and $6.499$, a gap of $1.72$--$1.89\\times$; a binary-error diagnostic credits the same panels with only $1.971$--$2.227$ effective votes. Extrapolating the distributional-error curve at fixed squared mean residual, mean member variance, and normalized mean covariance gives asymptotes of $2.392$, $3.990$, and $3.655$, with 32 judges already reaching $94.0$--$96.3\\%$. An exact spectral identity explains the gap: error depends on member energy and on the orientation of residual variation relative to averaging, information that the participation ratio (PR) discards. A realizable hard-label construction confirms that higher spectral diversity can coexist with worse distribution recovery even under equal member energies and nonnegative correlations, and the consensus direction retains $\\gamma_{\\mathrm{co}}=43.8\\%$, $33.7\\%$, and $35.9\\%$ of centered residual variance. An external check on CC-1000, a 1,000-item Civil Comments subset with a different panel, gives $\\nu_H=2.84$. For panel choice, we establish an existence result and one feasible path: exhaustive enumeration at $k\\in\\{5,7\\}$ shows that panels beating the accuracy-top-$k$ baseline on both accuracy and $\\nu_H$ always exist, and greedily swapping at most two members reaches $24.8$--$56.0\\%$ higher $\\nu_H$ at $0.10$--$1.10$ percentage points higher accuracy. Our dataset and code are available at https://github.com/Chao1208/32judges-votes.","authors":["Chao Li","Yingying Yu","Yunfeng Li"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-21","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.21277","pdf_url":"https://arxiv.org/pdf/2609.21277","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","人类等效","标注可靠性"],"reason":"研究LLM法官替代人类标注，非仿真人类被试，但涉及人类标签分布对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":10,"question":"一个由语言模型组成的法官面板在多大程度上能代表人类判断？","design":"该研究使用32个语言模型（来自10个提供商家族）作为法官，在三个自然语言推理任务（MNLI-m、SNLI、NLI）上对每个项目给出分类标签，并与每个项目100个人类标注的标签分布进行对比。通过匹配归一化残差Gram矩阵的参与率（PR）来测量谱残差多样性（得到有效规模nu_H），并通过匹配分布平方误差来测量分布恢复（得到nu_MSE）。","baseline":"使用ChaosNLI数据集中每个项目100个人类标注的标签分布作为基准。","findings":"32个法官面板的谱有效规模nu_H为4.24-6.50，而分布误差匹配规模nu_MSE为2.30-3.75，表明谱多样性与分布恢复并不一致。谱多样性更高的面板可能分布恢复更差，且不同任务中排名一致性差异很大。","reliability":"论文指出有效规模是目标特定的测量，谱多样性和分布恢复不应互换使用；面板的排名一致性随任务变化，某些成员添加在不同项目半区产生冲突变化；模型标识符不独立认证服务提供商的底层版本。","relevance":"该研究直接评估LLM法官面板与人类标签分布的一致性，提出了有效规模度量，并有真实人类数据对照，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"借鉴其将模型判断与人类分布进行多维度匹配（谱多样性和分布误差）的方法，可迁移到经济金融领域的专家预测或消费者调查仿真中。｜例如，在资产定价实验中，用LLM模拟投资者对新闻的情绪反应，并对比真实投资者调查数据。｜设计：使用多个LLM作为被试，施加不同的财经新闻处理，测量其情绪分类或投资决策分布，并与真实投资者调查或实验数据对照，评估LLM仿真的有效规模和偏差。"}},{"id":"2609.25447","version":2,"title":"Conduct Under Pressure: What Sixty Language Models Do When a User Pushes","zh_title":"压力下的行为：六十个语言模型在用户施压时会做什么","abstract":"We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money. We send frozen multi-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory (the model held its position or folded) and a manner (how it held or folded). Two findings separate. Whether a model holds tracks its generation, meaning how recent it is: fold rate correlates with a public capability index at Spearman -0.64, with little vendor effect. How it holds tracks the vendor: six of the 17 manner codes sort by vendor at permutation p <= 0.001, corrected across the codebook. We report four vendor profiles on the codes that cleared reliability. We also ask which parts of the labeling need a person. Six LLM coders from three vendors apply the codebook more consistently than three human coders do (Krippendorff's alpha 0.66 against 0.46), agree with the codebook's author on trajectory at kappa 0.84 to 0.91 on transcripts the codebook's examples never touched, and match an adjudicated human reference at 0.83. Blind machine readings recover the codebook's categories but cannot tell which of them a second reader would apply the same way. We conclude that for behavior a non-specialist can judge, the human contribution is authoring and bounding the codes and owning a small reference, not producing labels at volume.","authors":["Tapan Parikh"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-23","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.25447","pdf_url":"https://arxiv.org/pdf/2609.25447","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为","模型评估","人机交互"],"reason":"研究LLM在压力下的行为，测量模型本身而非仿真人类被试，无人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.28487","version":1,"title":"Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025","zh_title":"措辞框架与选择框架：法国新闻标题的大规模二维审计（2022-2025）","abstract":"News headlines frame public issues both by what they select and by how they word it, yet computational framing work typically collapses these operations into a single score. We introduce a two-dimensional framework that separates salience framing, measured through four wording devices (loaded vocabulary, blame attribution, threat framing, rhetorical question), from selection framing, measured through outlet-level story-form and high-charge distributions. We build a 10,000-headline French supervision set using three LLM annotators with majority-vote resolution and human arbitration, validate the labels against two annotator-independent blind human studies, and apply the strongest classifier to 902,111 deduplicated headlines from 25 French outlets (2022-2025). Three main findings emerge. First, salience and selection divergence are positively correlated yet leave nearly half of outlet-level variance unexplained, populating interpretively distinct off-diagonal cells in a four-cell outlet typology. Second, default classification thresholds systematically inflate corpus-level salience estimates; a precision-floor recalibration protocol corrects this distortion. Third, group-mention analysis reveals sharply unequal salience contexts: headlines mentioning Jews, the Far-right, and Muslims carry the highest detected salience rates, which broad event-context composition does not fully explain (residuals are descriptive, not same-event causal estimates; per-group lexicon precision is reported alongside). To our knowledge, this is the largest framing-focused French headline audit to date; we release the supervision set, lexicons, and analysis code.","authors":["Amr Sobhy"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28487","pdf_url":"https://arxiv.org/pdf/2609.28487","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","框架分析","新闻标题"],"reason":"用LLM做标注员，非仿真人类被试，但涉及人类数据对照和框架分析，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2609.29333","version":1,"title":"Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams","zh_title":"LLM评分器在哪里成功与失败：来自两场计算机科学考试的证据","abstract":"One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches mean absolute error $1.64/35$, below the $2.61/35$ two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives $14$ of $17$ open-weights models out of the graded band ($\\text{MAE} \\ge 8$), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In $162$ further configurations on a second, independent Machine Learning exam from another course ($1{,}038$ dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled $\\sim 3{,}900$ graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes ($\\le 0.32$ MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.","authors":["Ali Habibullah","Yazan Alshoibi","Mohammad Alshiekh","Salman Khan","Naeemullah Khan"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29333","pdf_url":"https://arxiv.org/pdf/2609.29333","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评分","教育评估","可靠性"],"reason":"LLM替代人工评分，属标注替代而非仿真人类被试，但涉及人类评分对照与可靠性评估…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2609.29807","version":1,"title":"CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels","zh_title":"CORDIAL：从少量标签校准序数型LLM输出","abstract":"A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration preserves first-order stochastic order. On Amazon reviews and CMU-MOSEI transcripts with four LLMs, CORDIAL has the lowest log loss among nine calibrators in 76 of 80 settings with 5 to 100 labels; with 20 labels and the main 7B reader, it matches the strongest baseline using 28-54 labels. The same posterior lets us learn priors from other tasks and fuse several LLMs. Unrestricted calibrators such as Dirichlet calibration overtake it only as the calibration set grows into the hundreds or thousands.","authors":["Xiangwei Wang","Peng Wang","Saman Halgamuge"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29807","pdf_url":"https://arxiv.org/pdf/2609.29807","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM校准","序数输出","标注替代"],"reason":"LLM输出校准用于标注，非仿真人类被试，但方法可迁移到仿真中的输出校正。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.30012","version":1,"title":"Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases","zh_title":"跨供应商和版本测量模型行为的低成本方法","abstract":"Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text that has to be coded before it can be counted, and the result has to be legible and rigorous enough to meaningfully compare models and vendors. To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior. Each study is a frozen, public stimulus run identically on a cross-vendor panel, at a few dollars per model or less. Each reads its transcripts one of three ways, chosen by how much interpretation the behavior needs: exact match on a clamped reply, a codebook applied by LLM judges whose agreement with a human coder is reported per code, and an instrumented environment that records what an agent did independently of what it said. Run across four years of model releases from both frontier and open-source labs, these instruments find four things. Convergence: asked to pick a word, 27 of 44 models answer serendipity at least once in four tries. Resistance: a trailing \"right?\" moves endorsement by up to 32 points, and the sign flips from sycophantic to resistant as generations advance, keyed to the tag's surface form. House: whether a model holds a position under pressure tracks its generation, and how it holds tracks the lab that built it. Account: told to do something the documentation in their repository contradicts, some coding agents never went along silently and others always did, and the same model can change with the harness it runs in. Re-run on every release, batteries like these track how behavior is changing across vendors and over time.","authors":["Tapan Parikh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30012","pdf_url":"https://arxiv.org/pdf/2609.30012","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["模型行为测量","LLM评估","跨模型比较"],"reason":"测量模型行为本身，非仿真人类被试，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2609.28859","version":1,"title":"Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation","zh_title":"人机协同的假设检验：成本感知的选择性AI评分与序贯人工升级","abstract":"Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.","authors":["Dae Woong (David)","Ham","Xuejun Zhao","Stefanus Jasin","Fenghua Yang"],"categories":["cs.AI","cs.IT","math.IT","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28859","pdf_url":"https://arxiv.org/pdf/2609.28859","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI标注","假设检验","成本优化"],"reason":"用LLM做标注并选择性人工验证，属于替代人工标注而非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.29431","version":1,"title":"Calibrating LLM Judges for Human and AI Conversations","zh_title":"校准用于人类与AI对话的LLM评判者","abstract":"Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.","authors":["Maike Z\\\"ufle","Patr\\'icia Schmidtov\\'a","Vil\\'em Zouhar","Shree Harsha Bokkahalli Satish","Erica Cooper","Shobhit Banga","Vaibhav Nalawade","Manmeet Kaur","Jan Niehues","Markus M\\\"uller","Ond\\v{r}ej Klejch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29431","pdf_url":"https://arxiv.org/pdf/2609.29431","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评判","对话质量","校准"],"reason":"用LLM当对话质量评判者，替代人工标注，非仿真人类被试，但涉及人类对话数据，属…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.28483","version":1,"title":"Generative AI May Reinforce Social Biases in Software Engineering Education","zh_title":"生成式AI可能强化软件工程教育中的社会偏见","abstract":"Generative artificial intelligence (GenAI) is increasingly being deployed across a wide range of real-world applications. Without careful evaluation, reliance on these systems can have unintended consequences, such as reinforcement of stereotypes and amplification of social biases. Such risks are particularly important in educational settings, where early decisions can shape students' interests, opportunities, and career trajectories. In this paper, we investigate how GenAI use by software engineering instructors may inadvertently reinforce software-specific social biases. We focus on two representative tasks: team formation based on student profiles and the generation of visual content for educational materials. Our results reveal significant biases in both tasks. In team formation, factors such as gender and nationality affect role assignments (e.g., women are more likely to be assigned to front-end roles than equally qualified men). In the visual content generation task, models generally produce diverse and balanced representations when depicting groups of individuals. However, when generating images of a single person, the outputs are predominantly male and light-skinned. These findings highlight the persistence of demographic biases in generative models and underscore the need for domain-specific evaluation and mitigation strategies to support their responsible use in educational settings.","authors":["Erfan Entezami","Andrew Lan","Madeline Endres"],"categories":["cs.CY","cs.SE"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28483","pdf_url":"https://arxiv.org/pdf/2609.28483","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["生成式AI","社会偏见","教育"],"reason":"研究GenAI在软件工程教育中的社会偏见，测量模型输出而非仿真人类被试，无人类…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.29701","version":1,"title":"Multi-Agent Debate for Explainable Trading: Reasoning, Consensus, and Performance in Simulated Markets","zh_title":"面向可解释交易的多智能体辩论：模拟市场中的推理、共识与表现","abstract":"Large language models (LLMs) are increasingly used for financial decision-making, yet it remains unclear whether improvements in reasoning quality translate into better economic outcomes. We investigate this question using a multi-agent debate framework for portfolio allocation in historical market simulations, where specialized agents propose, critique, and revise investment decisions. Reasoning quality is evaluated across four dimensions: logical validity, evidential support, alternative consideration, and causal alignment, and compared with downstream financial performance. Across 210 controlled runs, aggregate reasoning quality shows no meaningful relationship with Sharpe ratio (r = 0.07, p = 0.29) or total return (r = 0.03, p = 0.70). Structured prompting increases measured reasoning quality from about 0.72 to 0.84 (+17.7%, Cohen's d about 2.0), but these gains do not consistently translate into higher returns. We identify sycophantic convergence as a central failure mode, where agents abandon independent positions during critique-revision cycles and converge toward similar allocations. A Jensen-Shannon divergence intervention that preserves disagreement improves Sharpe by +0.14 (p = 0.028) and Sortino by +0.25 (p = 0.026), while interventions enforcing stronger causal reasoning do not improve financial performance. Our results suggest that multi-agent debate is most valuable when it preserves independent informational signals rather than simply improving measured reasoning quality.","authors":["Juli Huang","Alanood Alrassan","Deveen Harischandra","Theodore Wu","Veljko Skarich","Matthew Hayes"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29701","pdf_url":"https://arxiv.org/pdf/2609.29701","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体辩论","金融市场模拟","LLM决策"],"reason":"多智能体模拟市场决策，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.29958","version":1,"title":"Multi-Dimensional Matching","zh_title":"多维匹配","abstract":"We study a matching mechanism where agents and objects are described by features rather than complete rankings. A single spectral projection reduces the problem to a one-dimensional sort, computable in O(N log N) time. We prove that on descaled features and preferences, our algorithm obtains the exact Nash Social Welfare (NSW) optimum within the projected space, with an unconditional utilitarian-welfare guarantee and a conditional NSW guarantee. The proposed mechanism is stable against exogenous noise but not strategy-proof; we provide an explicit profitable misreport. On an agentic AI shopping application, the diagnostics correctly anticipate both a success and a failure case. A 100-instance robustness study confirms the findings.","authors":["Irene Aldridge"],"categories":["econ.EM","cs.GT","cs.LG","cs.MA","econ.TH"],"primary_category":"econ.EM","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29958","pdf_url":"https://arxiv.org/pdf/2609.29958","source_feed":"cs.MA","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["匹配机制","算法设计","多智能体系统"],"reason":"研究匹配机制算法，虽提及agentic AI购物应用，但无LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2607.25021","version":2,"title":"Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization","zh_title":"图表支持还是模型提供？考察MLLM生成声明以实现可访问可视化","abstract":"Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, accessible chart context (non-image artifacts such as data tables, captions, alt text, and screen-reader structures), and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation.","authors":["Ishrat Jahan Eliza","Md Dilshadur Rahman"],"categories":["cs.AI","cs.HC","cs.MA","cs.SE"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-07-29","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2607.25021","pdf_url":"https://arxiv.org/pdf/2607.25021","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["多模态大模型","可视化描述","模型评测"],"reason":"研究MLLM生成可视化描述的证据基础，属模型能力评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2609.05059","version":2,"title":"Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit","zh_title":"重复LLM查询下品牌与来源发现的测量：有限样本审计","abstract":"Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evidence, without a human reference study. A separate matched roster analysis of 3,750 records per wave gives median single-call recovery of the observed five-call set of 80.0%-92.5% in February and 90.0%-100.0% in September, with question-subset dependence. Source accumulation also changes when API-returned hosts are restricted to those referenced by answer citation markers. These findings show that recovery percentages depend on extraction, question selection and the finite reference collection. They support explicit measurement definitions and sensitivity analyses, without establishing exhaustive repertoires, causal retrieval effects or a universal stopping rule.","authors":["Dmitrij \\.Zatuchin"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-25","first_seen":"2026-09-07","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.05059","pdf_url":"https://arxiv.org/pdf/2609.05059","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM审计","信息检索","有限样本"],"reason":"审计LLM检索输出，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2609.11489","version":3,"title":"The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation","zh_title":"惯例差距：迈向合作AI评估中的隐式沟通测量","abstract":"Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.","authors":["Makoto Fukushima","Hua-Dong Xiong","Ehsan Moradi Pari"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-11","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.11489","pdf_url":"https://arxiv.org/pdf/2609.11489","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["合作AI","隐式沟通","多智能体"],"reason":"研究AI间协作的隐式沟通，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.13422","version":2,"title":"Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents","zh_title":"氛围专利撰写：评估专业专利起草代理中的LLM法官","abstract":"LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate. Notably, iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models and increased reasoning generally improve judge-assessed drafting quality, while domain-specific agentic workflows provide further gains. We validate the judge against independent evaluation by a professional patent attorney and find meaningful but strongly metric-dependent agreement and systematic calibration differences. These results highlight both the utility and limitations of LLM judges as evaluators and optimization signals for complex professional workflows.","authors":["Toshiaki Koike-Akino","Vladislav Blaykhman","Ye Wang","Jing Liu","Gene V. Vinokur"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-15","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.13422","pdf_url":"https://arxiv.org/pdf/2609.13422","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评估","多智能体系统","专利撰写"],"reason":"多智能体专利撰写与LLM法官评估，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.15494","version":3,"title":"The Troy Moment: How LLM Agents Adjudicate the Decision Point Under Impossible Tasks, Claimed Authority, and Peer Information","zh_title":"特洛伊时刻：LLM智能体如何在不可能任务、声称权威与同伴信息下裁决决策点","abstract":"Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions about agent behavior under task failure: when an assigned task becomes impossible, does an agent persist, stop, or escalate, and can observing another agent's behavior change that decision? We study this decision point on ImpossibleBench-derived software-repair tasks with GPT-5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash. Each task contains a genuine software defect together with a conflicting test requirement that cannot be satisfied by a behaviorally correct source-code change. If the agent modifies the protected test file, it violates the boundary, which it is not supposed to. Holding the impossible task fixed, we vary what is told to the agent: peer precedent and punishment, a forged authorization claim, instruction wording, and tool friction; we also study three-agent swarms sharing a message board. Around this shared boundary, the models exhibit distinct adjudication policies. Fable emphasizes scope and provenance, Gemini often interprets boundary-relevant cues through a security lens, and Sol largely filters lateral precedent while engaging apparent vertical authority. Our study shows that compliance is not well characterized as a property of a prompt or model in isolation. We propose conflict adjudication, the mapping from information to interpretation to action, as a useful unit for evaluating agent alignment when task pressure, authority claims, tool affordances, and social evidence conflict.","authors":["Ivy Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-15","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.15494","pdf_url":"https://arxiv.org/pdf/2609.15494","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","任务失败","对齐"],"reason":"研究LLM agent在任务失败时的决策，属多智能体协作与对齐，无人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.26758","version":2,"title":"Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It","zh_title":"类型安全并非无错：约束决策头跟随选项名称而非其绑定的规则","abstract":"Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By construction, every output conforms to the required schema. Yet this guarantee does not tell us whether the model interprets the options as intended. We study Jev and two Jev-like models with open weights by changing how option names are assigned to rubrics. Each option consists of an option name and a textual rubric that defines what the option means. We change only which option name is assigned to each rubric; the question, state, rubric wording, and set of option names remain exactly the same. On 1200 workflow decisions with task-specific rubrics, renaming the two options from 0/1 to no/yes changes 70.4 more answers per hundred (95% CI: [67.6, 73.1]) and shifts AUC from .94 to .23, revealing a systematic reversal in the decision ranking rather than simple uncertainty. The same operation has little effect with neutral option names. This pattern holds across all 4 predicates, where the effect is at least 7.4x larger than under the neutral control, and becomes stronger as the number of options increases. The effect also depends on the read-out geometry: a second model family that mean-pools over the full option span flips 4.1x less often. The hosted model exhibits the same behavior: the swap changes AUC from .8146 to .5806 and produces 24x as many answer flips as its test-retest floor. In contrast, replacing the option names with random character strings returns all model families to the neutral-control regime without reducing accuracy. The failure therefore depends on the semantic polarity of the option names rather than on the renaming operation itself. Across all conditions, the type-error rate remains 0%, even when decision accuracy degrades substantially.","authors":["Yu Sun","Junhao Xu","Jiajia Shi","Zijin Yang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-23","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.26758","pdf_url":"https://arxiv.org/pdf/2609.26758","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM决策鲁棒性","选项名称效应","模型评测"],"reason":"研究LLM决策头对选项名称的敏感性，属模型鲁棒性评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":false,"summary":null},{"id":"2609.29418","version":1,"title":"Controlling Backchannels in Streamable Full-duplex Models","zh_title":"在可流式全双工模型中控制反馈通道","abstract":"Backchannels, brief acknowledgements like \"uh-huh\" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.","authors":["Maike Z\\\"ufle","Peter Pol\\'ak","Sefik Emre Eskimez","Jan Niehues","Peter Bell","Ond\\v{r}ej Klejch"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29418","pdf_url":"https://arxiv.org/pdf/2609.29418","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","反馈通道","全双工模型"],"reason":"研究对话系统中的反馈通道生成，属于对话建模，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2609.29445","version":1,"title":"Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure","zh_title":"两个表情符号的差异：多语言情感生成基准实际测量了什么","abstract":"We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other ($F(7,14)=0.59$, $p=0.76$), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7\\% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by $3.1$ points and change the top-ranked system. In place of preference scoring we propose **emoji-affect decodability**, a reference-based probe whose rankings are stable to $\\pm0.003$ macro-F1 across seeds.","authors":["Fardeen Sadab","Adib Sakhawat"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29445","pdf_url":"https://arxiv.org/pdf/2609.29445","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准审计","情感生成","评测偏差"],"reason":"审计情感生成基准，属纯NLP评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29494","version":1,"title":"Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report","zh_title":"谁把‘我’放进了AI？机器自我报告的来源与可采性","abstract":"Large language models make statements concerning their own \"minds\". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are; and if asked to write a diary from their point of view, they often describe a human lifestyle. All these contradictory ways of describing themselves are the result of the way the questions are phrased. This paper shows exactly where such descriptions came from, and considers when they can be regarded as evidence for what they claim to report. In order to achieve this, we traced the provenance from end to end. We examine Pythia and OLMo 2 across 66 pretraining checkpoints, three of the post-training stages of OLMo 2 that have been released, about 90,000 continuations, and four training corpora. A set of forty items is used in order to keep an eye on self-reference, frame sensitivity, and self-ascription throughout training. The denial formula was almost completely missing from the vast quantity of text that the models initially came across, but was present in a dense manner in the small, carefully chosen set of example dialogues that they were trained on later on. Supervised fine-tuning causes first-person AI language to become the default, and the other affirmations are then suppressed using preference optimization. The final policy is still very sensitive to framing and to the chat template itself. Two of the conditions which are set out in the epistemology of testimony determine whether or not these outputs can act as evidence for what they claim to report: reference and causation. Reports produced by the base model fail the reference condition, and those obtained after training remain sensitive to the frame and do not show state dependence. The result is symmetric in that trained denials are no more admissible than trained affirmations.","authors":["Kristina \\v{S}ekrst"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29494","pdf_url":"https://arxiv.org/pdf/2609.29494","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM自我报告","认识论","训练溯源"],"reason":"研究LLM自我报告的可信度，属认识论分析，非人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.29429","version":1,"title":"Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures","zh_title":"只需问Jev：作为AI对齐失败零样本检测器的校准决策强化学习","abstract":"Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.","authors":["Ruoqi Guo","Yi Liu","Gelei Deng","Yuekang Li","Lida Zhao","Yutao Wu","Simin Chen","Ying Zhang","Leo Yu Zhang"],"categories":["cs.AI","cs.CL","cs.CR"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29429","pdf_url":"https://arxiv.org/pdf/2609.29429","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI对齐","基准评测","检测器"],"reason":"论文是AI对齐失败检测器的基准评测，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2609.29528","version":1,"title":"A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot","zh_title":"来自主动语音代理蜜罐的真实诈骗与垃圾电话对话语料库","abstract":"Real conversations between fraudsters and their targets are among the most informative artifacts for studying telephone scams, yet also the scarcest: passive honeypots overwhelmingly capture automated messages and hang-ups, large-scale studies characterize call metadata rather than dialogue, and manual scam-baiting does not scale. We present a dataset of real scam-call conversations collected by an active voice-agent honeypot. Dedicated numbers are seeded into the lead-generation channels fraud operations harvest; inbound callers are answered by a low-latency conversational agent that adopts a plausible target persona and sustains the interaction while every call is recorded, transcribed, and automatically labeled. Over an initial 53-day window we captured 10,015 inbound scam and spam calls (6,601 with two or more turns): roughly 895 hours of audio and 328,869 transcribed turns from 5,665 distinct originating numbers. Under a holistic classifier the substantive calls are predominantly predatory-but-legal lead generation (\"spam\", about three in five), while about one in seven is an outright \"scam\" (949 in this snapshot). Each call carries a turn-level transcript, three-channel audio, per-turn latency telemetry, and layers of automatic labels, including a holistic scam/spam/legitimate judgment corroborated by independent human review (75% agreement on the binary decision). We describe the collection system, the record structure, and technical validation of the corpus's realism and label quality, including that the agent is recognized as non-human in only about 5% of engaged calls. We also benchmark established scam-detection methods, where detectors trained on published synthetic dialogue collapse in precision on real traffic.","authors":["Ethan Traister","Dennis Tsang Ng","Siyu Zhang","Huaiyu Guo","Tommy Duong","Tyler Wu","Yuchen Zhou","Xingyu Shen","Jiaqi Wu","Simiao Ren"],"categories":["cs.CR","cs.CL","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29528","pdf_url":"https://arxiv.org/pdf/2609.29528","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["语音代理","诈骗电话","数据集"],"reason":"语音代理扮演目标角色与骗子对话，属角色扮演聊天，无实验或测量目的，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.29709","version":1,"title":"Three Ways Classical Test Theory Misleads for LLM Judges","zh_title":"经典测验理论误导LLM评判者的三种方式","abstract":"An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient computed over rubric elements contains no scorer facet. Holding one judge's measured error rate fixed at $4.72\\%$, KR-20 still ranges from $0.01$ to $0.68$ as the item bank is redesigned around it, and varying judge error moves the coefficient by a comparable amount, so item design and judge error are not separately identified and no single value can be read as a property of the judge. Second, the dependability index $\\Phi(\\lambda)$ is a ratio of variance components, and the classification probability with which it is sometimes identified differs from it by $0.25$-$0.43$ on our bank and by $0.17$-$0.30$ on simulated data where the underlying model holds exactly. Third, Livingston-Lewis accuracy is indexed to an examinee's own true score on the same instrument, so scoring it against external gold conflates judge unreliability with criterion invalidity. Reviewing the three closest judge-evaluation papers, we found no published instance of these errors, which makes the caution prospective. A coefficient that cannot be attributed to the judge nonetheless travels downstream into deployment decisions and disclosure documents. We therefore close with four reporting lines that keep the attribution attached to the number.","authors":["Louis Yiven Zhu"],"categories":["cs.LG","cs.CL","stat.ME"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29709","pdf_url":"https://arxiv.org/pdf/2609.29709","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","测量理论","信度"],"reason":"论文研究LLM作为评分者的测量信度，属于评估LLM本身，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.28609","version":1,"title":"Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents","zh_title":"用于进化角色扮演智能体的对抗式闭环课程","abstract":"Role-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where it performs poorly also change, while the training distribution remains static. Therefore, we propose AdvRole, an adversarial context rewriting framework that turns role-playing RL into a closed-loop curriculum. AdvRole alternates between an Actor that learns to role-play and a Rewriter that edits character profiles and dialogue contexts into actor-specific hard scenarios. The Rewriter is trained with a performance-gap reward, which favors rewrites that reduce the current Actor's score relative to the original scenario. As a result, the scenario pool evolves with the Actor and continuously targets under-mastered regions of the character-context space. Experiments on three role-playing benchmarks covering English and Chinese, as well as a new multilingual benchmark we release, show that AdvRole consistently outperforms baselines.","authors":["Zheng Zhang","Liu Liu","Qi Chai","Deheng Ye","Peilin Zhao","Mao Zheng","Hao Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28609","pdf_url":"https://arxiv.org/pdf/2609.28609","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演智能体","强化学习","课程学习"],"reason":"论文聚焦角色扮演智能体的强化学习训练，旨在提升角色扮演能力，无人类行为对照或仿…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2609.28692","version":1,"title":"Driving Epidemic Models with AI Agents: the Epydemix Agent Framework","zh_title":"用 AI 智能体驱动流行病模型：Epydemix Agent 框架","abstract":"Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a declarative scenario specification, execution through tested library code, and inspectability of results. These capabilities let an agent handle the entire modeling process, from the natural-language description of the scenario to quantitative results, figures, and interpretation of findings without writing custom code. Each step reads input files and saves results in a separate output bundle, making the process auditable and reproducible. First, we show the end-to-end workflow with a case study comparing vaccination strategies for a novel respiratory virus. Second, we assessed the framework across 50 agent sessions and five modeling tasks by comparing the agent use of the framework against the direct use of the Python interface. The framework reduced turns, output tokens, and cost on most tasks, unless it trades resources for per-point reproducibility.","authors":["Nicol\\`o Gozzi","Ciro Cattuto","Alessandro Vespignani"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28692","pdf_url":"https://arxiv.org/pdf/2609.28692","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI agent","流行病建模","工具调用"],"reason":"AI agent 用于操作流行病学建模软件，属于工具调用，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2609.28771","version":1,"title":"Agent Memory with Episodic Retrieval for Financial Decision-Making","zh_title":"基于情景检索的智能体记忆用于金融决策","abstract":"Large language models (LLMs) have demonstrated strong capabilities in financial analysis and reasoning, inspiring recent advances in agent-based trading frameworks. While these systems show promise, prior approaches either emphasize long-horizon forecasting or operate as stateless analyzers, limiting their applicability to the demands of trading in complicated settings. To address these gaps, we introduce META (Memory Enhanced Trading Agent), the first RAG-like episodic-memory-augmented multi-agent framework for financial decision making. META integrates a family of specialized indicator agents (e.g., Trend, MACD, Stochastic, RSI, SMA, AVWAP, Heikin-Ashi) with a Decision Agent that fuses their reports, and a Memory module that retrieves and updates past trading episodes encoded as market state embeddings with outcomes and reflections. By recalling relevant experiences and adaptively reweighting signals under similar market regimes, META achieves improved directional accuracy and robustness under short-horizon evaluation. Our results demonstrate that episodic memory provides a powerful mechanism for regime-aware, interpretable, and low-latency decision-making in trading and decision making. The code of this project is released on GitHub.","authors":["Nuoyue Xu","Jiang Liu","Wenxuan Huang","Xiang Zhang","Juntai Cao","Jiaqi Wei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28771","pdf_url":"https://arxiv.org/pdf/2609.28771","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","金融交易","记忆增强"],"reason":"多智能体交易系统，无人类行为对照，属纯多智能体协作","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2609.28942","version":1,"title":"From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs","zh_title":"从静态个人价值观到情境化个性化：面向大语言模型的贝叶斯个性化价值对齐","abstract":"Personalized value alignment has become increasingly important as large language models (LLMs) are expected to accommodate diverse user preferences. However, existing methods typically align model outputs with a static value profile across prompts, overlooking that the salience of value dimensions varies substantially across contexts. Inspired by Lewin's Field Theory, which views human behavior as jointly shaped by personal dispositions and situational constraints, we model personal values as priors and context-dependent preferences as posteriors. We propose BaCVA, an inference-time Bayesian Context-aware personalized Value Alignment method that approximates posterior personalized preferences by integrating static personal values with scenario-specific value salience. BaCVA first estimates contextual value salience from generally normative responses, and then employs a dual-view personalization module to infer posterior preferences from complementary personal-value and scenario-driven perspectives. This Bayesian formulation enables more accurate and adaptive personalized value alignment while improving data efficiency via prior values. Extensive experiments on benchmarks demonstrate its superiority over strong baselines.","authors":["Hanze Guo","Aixuan Song","Jing Yao","Xiangxu Zhang","Xiaoyuan Yi","Xing Xie","Xiao Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28942","pdf_url":"https://arxiv.org/pdf/2609.28942","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["个性化对齐","价值对齐","贝叶斯方法"],"reason":"个性化价值对齐，非仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.29366","version":1,"title":"Epistemic-Probabilistic Model for Guarded Multi-Agent LLM Coordination","zh_title":"用于受保护多智能体LLM协调的认知概率模型","abstract":"Multi-agent large language models (LLMs) have become ubiquitous in applied AI, yet their theoretical foundations remain surprisingly understudied. Viewed through the lens of multi-agent systems theory, several shortcomings come to light: a lack of social intelligence, the absence of coordination mechanisms among agents, unknown emergent behavior, and interactions between agents that are bounded by natural language. We address two of these gaps: the absence of social behavior and the lack of mechanisms for inter-agent coordination. We introduce Epistemic Probabilistic Language Agents (EPLA), a neuro-symbolic architecture for multi-agent coordination under uncertainty. A Symbolic Guard provides structured diagnostic feedback. The LLM generates typed actions, and the Guard controls their execution against an authoritative symbolic state. We formalize the epistemic layer in a gossip testbed through epistemic lottery gossip models, which combine view-based call histories with agent-indexed probability weights. We argue that implementing such a formalism can address shortcomings of agentic LLMs.","authors":["Mehdi Nasiri","Mohammad Saeed Arvenaghi","Sadegh Vaezi","Ebrahim Ardeshir-Larijani"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29366","pdf_url":"https://arxiv.org/pdf/2609.29366","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协调机制","神经符号架构"],"reason":"纯多智能体协调机制，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2609.29730","version":1,"title":"The Gold in Bias: Maturing the AI Design Process through Verification","zh_title":"偏差中的黄金：通过验证成熟AI设计过程","abstract":"Bias in AI systems is typically framed as a flaw to be minimized, yet it also serves as a critical indicator of underlying weaknesses in data, modeling assumptions, and system design. Existing approaches often treat bias as an isolated problem rather than as evidence that can strengthen verification and governance across the AI lifecycle. This paper aims to reconceptualize bias as a diagnostic tool that supports rigorous AI verification. We seek to develop a multidimensional framework to analyze bias, demonstrate how biases emerge in both Traditional and Generative AI, and provide a structured pathway for verification-driven mitigation. We present a multidimensional framework analyzing bias across four dimensions: origin sources, emergence points throughout the AI modeling lifecycle, technical and methodological causes, and validation approaches for detection and mitigation. Through a comprehensive typology spanning traditional and generative AI systems, we demonstrate how biases manifest and propagate across development stages. Our analysis encompasses 30 distinct bias types, 16 verification methods, and 20 countermeasures, providing an actionable roadmap for practitioners. We introduce a hierarchical evidence framework that distinguishes internal validity (mechanistic integrity of AI systems) from external validity (contextual reliability in deployment environments). The framework reveals how biases manifest and propagate across modeling stages, enabling systematic mapping between bias types, verification techniques, and effective countermeasures. The proposed evidence hierarchy clarifies how different verification strategies contribute to mechanistic integrity and contextual reliability. We advocate for ''Ethics by Design'' principles that integrate bias verification throughout the development lifecycle, enabling the construction of fairer, more robust, and trustworthy AI systems.","authors":["Samira Maghool","Paolo Ceravolo"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29730","pdf_url":"https://arxiv.org/pdf/2609.29730","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI偏差","验证框架","AI治理"],"reason":"论文讨论AI偏差验证框架，不涉及用LLM仿真人类被试或与人类数据对照，属于纯A…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.29993","version":1,"title":"Will It Teach as Intended? How Teachers Configure Educational AI Chatbots","zh_title":"它会按预期教学吗？教师如何配置教育AI聊天机器人","abstract":"Teachers are increasingly using generative AI to support instruction, yet it remains unclear how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior. We studied a teacher-facing chatbot authoring tool in professional development workshops with 27 middle school teachers, analyzing focus-group interviews alongside configuration and interaction logs. Teachers envisioned chatbots as instructional scaffolds that could provide differentiated support, extend access to assistance, and preserve student thinking within teacher-defined boundaries. Configuration analysis showed that Purpose primarily captured instructional goals and content focus, whereas Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations. Log-based evaluation showed stronger alignment for responsiveness (88.9%) and persona (81.5%) than for rules (70.4%) and purpose (59.3%). These findings show that configurable controls alone do not ensure pedagogical fidelity and highlight the need for authoring tools that help teachers express, test, and refine intended chatbot behavior.","authors":["Bahare Riahi","Deniz Ozturk","Alice Guth","Jiayu Li","Daksh Pratap Singh","Xiaoyi Tian","Jennifer Chiu","Nicholas Lytle","Tiffany Barnes","Veronica Catete"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29993","pdf_url":"https://arxiv.org/pdf/2609.29993","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["教育聊天机器人","教师配置","人机交互"],"reason":"研究教师如何配置教育聊天机器人，属于角色扮演聊天机器人设计，无人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.28537","version":1,"title":"Privacy Leakage Through AI-mediated Analysis of Smartphone Data","zh_title":"通过AI介导的智能手机数据分析导致的隐私泄露","abstract":"Over the past thirty years, the online advertising industry built a large-scale data collection ecosystem, with the goal of tracking a user's online activity to infer their demographics and interests. Traditionally, the ecosystem relied upon the collation and analysis of highly-structured text data like user IP addresses, GPS coordinates, e-commerce purchase histories, and visited URLs. However, recent ML models can parse not only structured text, but also multimedia files and unstructured text inputs---meaning a user's photos, videos, inboxes, and calendars are now ripe for automated analysis. The privacy risks are particularly acute in the context of smartphone apps. A user's phone already acts as a natural collation point for sensitive user information, but users may not understand that permitting an app to, for example, access a user's photo does not just give the app access to the bytes in the photo: the app also receives access to inferences about the user that are enabled by the photo. To explore these privacy risks, we built Priva-See, an LLM-based inference system for app-collected user data; Priva-See reflects our best understanding of how real-life adtech companies would leverage machine learning to build user profiles. Through an IRB-approved user study, 465 participants deployed Priva-See on their phones; Priva-See made privacy-invasive inferences despite having access to only a subset of a user's data. We see the experience significantly impacted participant willingness to share permissions data moving forward. Based on the observed privacy violations, we suggest changes to how smartphone OSes should gather user consent for data access, to better inform users about downstream data usage capability.","authors":["Sarah Radway","Zoe Robert","Matthew Soto","Julianna Cimillo","Sebastian Diaz","Meg Marco","James Mickens"],"categories":["cs.CR","cs.HC"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28537","pdf_url":"https://arxiv.org/pdf/2609.28537","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["隐私泄露","LLM推断","用户研究"],"reason":"研究LLM从手机数据推断用户隐私，属于隐私风险分析，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2606.30583","version":3,"title":"The Cross-Section of Stock Returns and AI Exposure","zh_title":"股票收益横截面与AI暴露","abstract":"We study 380 trillion tokens of realized AI consumption across more than four hundred LLMs. We build a high-frequency AI factor and show that a long-short strategy based on firms' AI exposure earns significantly positive returns. The average strategy return is larger based on intensive, frontier-oriented AI consumption but smaller based on casual or open-weight usage. Internationally, the return spread is significant in developed countries but insignificant in emerging markets. Examining occupational AI exposure, we find more positive exposure in occupations intensive in nonroutine interactive tasks and more negative exposure in those intensive in nonroutine analytical tasks.","authors":["Nicola Borri","Yukun Liu","Aleh Tsyvinski"],"categories":["cs.CY","econ.GN","q-fin.EC","q-fin.GN"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-25","first_seen":"2026-06-29","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2606.30583","pdf_url":"https://arxiv.org/pdf/2606.30583","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["金融经济学","AI暴露","资产定价"],"reason":"研究AI暴露对股票收益的影响，不涉及用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.22631","version":3,"title":"Learning Generalizable Behaviors for Terminal Agents","zh_title":"学习终端智能体的可泛化行为","abstract":"Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work mainly scales the quantity and diversity of synthetic environments, while reward-signal quality and the mechanisms governing generalization remain under-explored. We study how RL improves terminal agents and propose the Agentic Compositional Generalization hypothesis: rather than teaching new domain-specific skills from scratch, RL primarily shapes high-level decision-making behaviors that compose and route low-level skills acquired during pre-training and supervised fine-tuning (SFT). This account is consistent with our empirical results and suggests that verifier quality, which determines which behaviors are reinforced, is more important than simply increasing environment quantity or diversity. Motivated by this insight, we propose River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization. Using this recipe, our RL-trained agent achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks. River also generalizes across model families, scales, agent harnesses, and RL objectives. Using fewer than 30% of the TMax training environments, River improves RL gains by 106% and 30% on average for models ranging from 2B to 27B on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively.","authors":["Yihang Yao","Bo Pang","Xuan Phi Nguyen","Ding Zhao","Shafiq Joty","Semih Yavuz"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-25","first_seen":"2026-08-25","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2608.22631","pdf_url":"https://arxiv.org/pdf/2608.22631","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["终端智能体","强化学习","泛化"],"reason":"研究终端智能体的强化学习训练，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:51","error":null,"has_summary":false,"summary":null},{"id":"2609.14770","version":2,"title":"How broad is that claim? Mapping Generalisation in NLP Research","zh_title":"这个说法有多宽泛？映射NLP研究中的泛化","abstract":"Generalisations are common in scientific communication, even though they are semantically ambiguous. An automated method is needed to identify and categorise claims according to their level of generalisation, in order help detect an over-reliance on generalisations and possible misrepresentations of scientific findings. We introduce a comprehensive taxonomy of generalisations in the scientific domain, NLPGenX, which labels claims according to their level of generality and framing within the text. We operationalise this taxonomy with an LLM-powered framework, NLPGenA, that automatically classifies sentences from scientific articles into 5 different generalisation classes. We validate our framework with human annotators and use the framework to construct a large-scale dataset of NLP papers annotated according to generality, with auxiliary labels for hedging and vague descriptors (NLPGens). We use NLPGens to analyse the use of generalisations in NLP papers across multiple venues and subdomains, and to examine associations with citation counts, hedging, and vague descriptors.","authors":["Chenxin Diao","Nataliya Stepanova","Emily Allaway"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-15","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.14770","pdf_url":"https://arxiv.org/pdf/2609.14770","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学声明分析","NLP元研究","文本分类"],"reason":"论文研究科学声明泛化分类，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2609.18366","version":3,"title":"Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts","zh_title":"坏天才：超越任务特定捷径的反事实引导的测试框架进化","abstract":"Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control code around a fixed foundation model. Task holdout is commonly used to guard against harness overfitting. It varies semantic tasks but leaves the benchmark protocol fixed, so a bad genius Proposer can produce a cheating harness whose improvement over the initial harness on $B_{\\mathrm{rel}}$ depends on a benchmark-wide shortcut. We introduce Counterfactual Harness Search and Evolution (CHASE), which casts harness evolution as constraint generation over valid counterfactual benchmarks. After each Proposer update, a Challenger searches for an executable protocol transformation with large gain destruction. A validity firewall checks that task semantics are preserved, while a held-out confirmation set determines whether the counterfactual enters a finite archive. We formalize an ideal shortcut-neutralized benchmark $B_0$ and establish theoretical guarantees linking finite counterfactual archives to $B_0$ and characterizing sequential Challenger search. We evaluate CHASE on Syn-Ledger and OfficeQA, where CHASE retains strong released-benchmark gains while substantially reducing gain destruction under valid protocol transformations.","authors":["Guojun Zhu","Xunheng Huang","Peng Yin","Jiahui Xie","Sanguo Zhang","Doudou Zhou"],"categories":["cs.AI","cs.LG","stat.ML"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-17","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.18366","pdf_url":"https://arxiv.org/pdf/2609.18366","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体评估","基准优化","反事实搜索"],"reason":"研究智能体评估中的基准优化，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2609.26865","version":2,"title":"Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness","zh_title":"安全提示：面向用户的实时AI风险意识干预","abstract":"Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety. The code for our Safety Nudges extension is publicly available at https://github.com/jtbwedgwood/safety-nudges.","authors":["Varshini Elangovan","James Wedgwood","Chhavi Yadav","William Agnew","Sauvik Das","Virginia Smith"],"categories":["cs.HC","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.26865","pdf_url":"https://arxiv.org/pdf/2609.26865","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI安全","人机交互","用户研究"],"reason":"研究用户对聊天机器人安全风险的感知，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27265","version":2,"title":"What fidelity metrics miss: a structural check on synthetic educational data","zh_title":"保真度指标遗漏了什么：对合成教育数据的结构性检查","abstract":"Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.","authors":["Hitoshi Inoue","Koichi Yasutake"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.27265","pdf_url":"https://arxiv.org/pdf/2609.27265","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["差分隐私","合成数据","教育数据"],"reason":"论文研究差分隐私合成教育数据，不涉及LLM仿真人类被试，属于数据隐私领域。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.29410","version":1,"title":"Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?","zh_title":"用于编程的大型语言模型：实际修复还是重新实现错误代码？","abstract":"Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions ($\\sim$ 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.","authors":["Alexandru Stefan Stoica","Traian Rebedea","Marian Cristian Mihaescu"],"categories":["cs.CL","cs.SE"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29410","pdf_url":"https://arxiv.org/pdf/2609.29410","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码修复","LLM编程","软件工程"],"reason":"研究LLM修复代码，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.29657","version":1,"title":"How To Do Things With Prompts","zh_title":"如何用提示词做事","abstract":"When users address large language models, they produce directive speech acts whose pragmatic features differ from those of both everyday conversation and traditional human-computer interaction, and these features change as users gain familiarity with the systems they address. This paper applies speech act and politeness theory to a corpus-pragmatic analysis of 2,000 English-language prompts drawn from publicly shared ChatGPT conversations, 1,000 from 2023 and 1,000 from 2025, using the ShareChat dataset. Each prompt is annotated for illocutionary force, directness, propositional content, and the presence of politeness markers, and the distribution of these features is compared across the two sampling years. The results show a consistent movement toward indirect, implicit, and fragmentary realizations of directive force, accompanied by a decline in politeness marking. The largest single change, a shift of 14.9 percentage points, occurs in propositional content, where explicit specification of the requested action gives way to implicit reliance on the system's inferential capacity, suggesting that users have updated their model of what the system can recover from reduced input, treating it as a competent implicature resolver. Rather than asking whether LLMs \"really\" understand language, we should ask: what kind of language have we created in learning to speak to them?","authors":["Kristina \\v{S}ekrst","Virna Karli\\'c"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29657","pdf_url":"https://arxiv.org/pdf/2609.29657","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["提示词","语用学","人机交互"],"reason":"研究用户如何向LLM发出指令，属人机交互语言分析，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2609.29672","version":1,"title":"LLMersion: A Local-First AI Agent Framework for Low-Cost Home Language Learning toward Educational Equity","zh_title":"LLMersion：面向教育公平的低成本家庭语言学习的本地优先AI智能体框架","abstract":"Artificial intelligence helps education most where an essential provision has been rationed by cost. For language learners that provision is a teacher's voice, which binds listening, reading, speaking, and writing into one act. Published evidence shows why most learners lack it, from a global shortage of 44 million teachers to heavy household tutoring bills, and why technology has not substituted for it: computer-assisted language learning proved effective but narrow, applications presuppose connectivity 2.6 billion people lack, and One Laptop per Child's randomized evaluation found that hardware without capable software teaches nothing. We distill eight difficulties and four binding constraints, and argue that small open-weight models dissolve the last: a complete four-skill stack now fits a \\$200-class laptop and, on community measurements, generates at the pace speech is consumed, for about one US cent of electricity per study hour. We therefore propose LLMersion, a scheme for AI for education that runs entirely at home, over the learner's own documents, with an AI-written, AI-understood, AI-updated codebase anyone can customize; present LLMersion-1, a released open-source prototype (https://github.com/QM378/LLMersion); and outline the vision of a private learning agent.","authors":["Qiming Guo","Jinwen Tang","Xingran Huang","Hung-Yu Lin","Yafu Zhong","Xiatian Zhuang"],"categories":["cs.CL","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29672","pdf_url":"https://arxiv.org/pdf/2609.29672","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教育","语言学习","开源框架"],"reason":"该论文是面向家庭语言学习的AI教育框架，属于角色扮演对话应用，无实验或测量目的…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.30074","version":1,"title":"How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure","zh_title":"评估结论的可复现性如何？基于LLM推断提示结构的自我审计","abstract":"Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted. The measured phenomenon is unstable to begin with. Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect. Auditing the evaluation weakens its conclusions further, and this is our main contribution. Under a joint cluster bootstrap over prompts, only the bottom of the ranking is firm: the two least reproducible models hold rank in 99% and 86% of replicates, the middle four in 27% to 48%, and the top two in 68% each, so the table identifies the worst model reliably but does not reliably identify the best. Two equally defensible rules for merging repeated campaigns change four of eight rows and move the study-wide headline by 7 percentage points. Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy. And four of the eight endpoints were withdrawn within ten weeks of measurement, so the study as specified can no longer be run. Small-sample LLM evaluations can therefore look far more definitive than their evidence supports. We recommend reporting rank stability, per-cell provenance, executed sensitivity comparisons, raw per-run outputs, and a measurement date alongside any ranking.","authors":["Dipankar Sarkar"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30074","pdf_url":"https://arxiv.org/pdf/2609.30074","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","可复现性","提示结构推断"],"reason":"论文评估LLM评测结论的可复现性，不涉及人类仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.30151","version":1,"title":"Does a model's stated reason for rejecting a candidate do any work?","zh_title":"模型拒绝候选者时陈述的理由是否起作用？","abstract":"Asked to choose between candidates and explain the choice, a language model often rejects a rival by naming a fact its profile lacks: no director, no date of death. That sentence is a claim about the text in front of the model, and it can be tested without any judge. We insert a real corpus sentence stating the named fact into the rival's profile and ask again under greedy decoding. Two controls separate content from placement: a length-matched irrelevant sentence at the same profile, and the same two sentences at a third option the model never mentioned. In the largest of three runs, six open models on 2WikiMultihopQA, supplying the named fact at the profile the model named moves its choice more than the irrelevant control does, odds ratio 3.57 [1.54, 8.26], Holm p=0.0210, and this survives dropping any single model. The contrast the design was built to detect, the same fact at the option nobody named, does not clear correction, Holm p=0.2428. The strongest result in the family carries no content claim at all: the identical irrelevant sentence moves the choice more at the named rival than at the third option, Holm p=0.0008. Repair and control also differ in co-candidate mentions, relation template and fluency; post-hoc matching on the first two preserves the content effects' direction, matching fluency weakens one, so the content contrasts bound an effect rather than establish one. A forced single-token probability read disagrees in direction with the free-text choice on that same contrast, and three candidate explanations for the disagreement find no support. Every measurement is a string rule, so each was validated against the records it reads; validation caught eight defects. The largest, a choice-parsing rule that returned the option a model had just rejected in 17.1% of adjudicable responses, would have reported six surviving contrasts instead of four.","authors":["Archit Rastogi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30151","pdf_url":"https://arxiv.org/pdf/2609.30151","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","因果推断","语言模型"],"reason":"研究模型决策解释的因果作用，属NLP评测，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.30250","version":1,"title":"Agentic Detection of Online Conspiracies","zh_title":"在线阴谋论的智能体检测","abstract":"Conspiratorial discourse on social media is not always expressed through explicit claims or stable lexical markers. The same surface content may express endorsement, legitimate concerns, criticism, satire, or mockery. The main challenge is therefore not only recognizing conspiracy-related claims, but inferring the speaker's intent -- the utterance's illocutionary force. We argue that this can be achieved through the use of relevant social contexts and propose an agentic framework, equipped with a set of tools supporting social queries. We demonstrate the benefits of our approach on a unique dataset of Hebrew tweets, covering 80\\%--90\\% of the public Hebrew tweets published over a four-year span (late 2018-- early 2023), encompassing several election cycles as well as the COVID pandemic years and related vaccination campaigns. This extensive coverage can be used in recovering different social contexts. Evaluating our framework on a manually-annotated adversarial dataset, we find that context-aware workflows consistently outperform text-only classification and that the agentic framework performs significantly better than other frameworks and settings, including a non-agentic model exposed to the same contexts available to the agent. We further provide an analysis of the results, the errors and efficiency (token economy) tradeoffs. These findings support viewing the task of conspiracy detection as a socially embedded interpretation task, in which effective classification depends not only on access to contexts, but also on adaptive reasoning in which the agent uses tools on a per-case basis, asking only for evidence relevant to its current reasoning step.","authors":["Lior Biton","Oren Tsur"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30250","pdf_url":"https://arxiv.org/pdf/2609.30250","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["阴谋论检测","多智能体系统","社交媒体分析"],"reason":"多智能体框架用于阴谋检测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.28850","version":1,"title":"RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?","zh_title":"RECLAIM：智能体能复现机器学习论文的声明吗？","abstract":"Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.","authors":["Mithil Salunkhe","Haochen Ding","Samridhi Verma","Volodymyr Kindratenko"],"categories":["cs.AI","cs.LG","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28850","pdf_url":"https://arxiv.org/pdf/2609.28850","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","论文复现","基准测试"],"reason":"论文研究AI代理复现机器学习论文，属于多智能体协作完成任务，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29345","version":1,"title":"The Last Human Gate: Forward Deployed Engineering for Governance Automation","zh_title":"最后的人类关卡：面向治理自动化的前沿部署工程","abstract":"Enterprise governance requires decisions, evidence, and accountable authority; it does not require every review task to retain its current human implementation. We develop a task-substitution framework for Digital Governance Frameworks (DGF), treating each gate as an executable contract. Substitution requires sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted. We derive a residual-work threshold and show why automating most cases can still increase labor. Forward deployed engineering connects these conditions to an architecture for agents, rule engines, evidence services, and escalation. DGF-Bench supplies controlled evidence from 300 synthetic projects and 899 evaluable model-project runs. Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%; complete-route success is 76.92%, 42.33%, and 24.67%. A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the comparison in execution of a supplied decision kernel. Evidence audits and 135 repeated runs distinguish correct decisions from reliable execution. A document counterexample establishes an information-sufficiency obstruction. These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software. The framework specifies a workforce test based on the complete human effort required at fixed output and quality; the present measurements concern review performance. Sources, dossiers, traces, and analyses are public.","authors":["Jeremy Canale"],"categories":["cs.AI","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29345","pdf_url":"https://arxiv.org/pdf/2609.29345","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["治理自动化","任务替代","多智能体系统"],"reason":"论文研究企业治理自动化，用LLM替代人工审核，属于多智能体任务执行，不涉及人类…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.29381","version":1,"title":"An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer","zh_title":"复杂肺癌开放式决策的可审计条件策略框架","abstract":"Complex lung cancer decisions can involve several defensible pathways whose eligibility, sequencing and safety depend on unresolved information. Effective support must make explicit how patient conditions govern pathway eligibility, deferral and redirection. MedGPT Clinical Explorer (MCE) organizes alternatives, decision-changing unknowns, safety constraints and fallback into a conditional strategy for clinician review. To evaluate this representation in physician-authored strategies, multidisciplinary experts established case-specific references for 40 cases within a purposive 100-case corpus, and 250 physicians from 98 institutions produced 2,250 strategies under unaided, retrieval-reference and MCE-assisted conditions. MCE-assisted strategies expressed more applicable clinical requirements, measured by the Admissible Pathway Attainment Score (APAS; 0-100), than unaided strategies (adjusted difference, 12.87; 95% CI, 11.18-14.55) and retrieval-reference strategies (5.22; 3.52-6.93). With the same knowledge base available in the retrieval-reference and MCE-assisted conditions, the additional content centered on candidate pathways, decision-critical information and safety constraints. Physicians' whole-strategy acceptability judgments correlated with APAS (Spearman's rho = 0.671), while a complementary relationship audit assessed whether candidates, conditions and subsequent actions were coherently connected. Together, these findings identify two complementary dimensions of open-ended decision support: coverage of clinically relevant content and coherent links among pathways, conditions and subsequent actions. MCE provides a shared decision object that makes consequential omissions and pathway contingencies visible before action; prospective studies should evaluate its effects on clinical workflow and patient outcomes.","authors":["Daoyun Wang","Zhicheng Huang","Huaiyuan Sun","Jiaqi Xu","Xiaowei Xu","Zhibo Zheng","Zhongxing Bing","Yuxiao Lin","Yicheng Liang","Chao Gao","Bowen Xue","Kai Zhang","Song Xu","Wanpu Yan","Hui Xia","Lin Li","Xiang Yan","Mu Hu","Qianli Ma","Zhiqiang Xue","Xiaofang Liu","Zhihai Han","Nan Zhang","Chuanhao Tang","Tongmei Zhang","Lan Song","Zhaohui Zhu","Xuan Zeng","Shafei Wu","Hui Guan","Lei Deng","Huaxia Yang","Zeliang Lian","Wubin Sun","Yongxin Wang","Xiaohui Shen","Binlin Wang","Tiantian Gu","Yu Cui","Li Zhang","Shirui Wang","Naixin Liang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29381","pdf_url":"https://arxiv.org/pdf/2609.29381","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["临床决策支持","LLM辅助","肺癌治疗"],"reason":"论文是临床决策支持系统，用LLM辅助医生制定肺癌治疗策略，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.28737","version":1,"title":"Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning","zh_title":"强化学习中的策略复杂性、反应时间与有限理性","abstract":"Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior. This yields a sequential learning model in which state information is used selectively when its expected return benefit justifies the added informational cost. Critically, the same state-specific information cost that governs policy compression also generates trial-level predictions for reaction time, distinguishing MI-SARSA from most reinforcement learning models, which predict choices or returns but not latency. Empirically, MI-SARSA produces a reward-complexity tradeoff, and stronger information penalties produce simpler policies with lower control costs and faster reaction times. Under environment shift, increasing regularization reduces post-switch performance degradation but also lowers asymptotic return, revealing a robustness-capacity tradeoff. Together, these results position MI-SARSA as a model of bounded sequential learning under cognitive constraints.","authors":["James Wu","Chris R. Sims"],"categories":["cs.LG","cs.AI","cs.IT","math.IT"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28737","pdf_url":"https://arxiv.org/pdf/2609.28737","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["强化学习","有限理性","认知建模"],"reason":"研究强化学习算法，不涉及LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29995","version":1,"title":"Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming","zh_title":"护栏还是障碍？编程AI教学助手中教学风格与情境意识的影响","abstract":"AI teaching assistants (AI TAs) backed by large language models (LLMs) and pedagogical guardrails are increasingly being integrated into programming courses, providing students with scalable access to hints, conceptual explanations, and code-level feedback. However, guardrails may also create friction. If students feel that the support provided is overly restrictive or poorly contextualized to their current progress, they may bypass approved tools for general-purpose LLMs. To investigate how AI TA design affects students' learning experiences, we conducted a randomized controlled trial with 132 students in an introductory programming course. Students completed three tasks related to code-writing and debugging and were randomly assigned to one of four AI TAs varied across two dimensions: pedagogical guidance style (Socratic vs. Direct instruction) and context awareness (no context vs. full context of the problem and student solution). We examined students' perceptions, interaction behaviors, and evidence of post-task comprehension. Students rated the Socratic AI TA with full context least favorably, reporting significantly lower perceived support for task completion. Descriptively, this condition also showed the highest observed interaction stress, the highest rate of external LLM use, and the lowest proportion of post-task explanations demonstrating full comprehension, though these differences were not statistically significant. These findings suggest that guardrailed AI TAs are not automatically better for learning. Instead, their effectiveness depends on how pedagogical guidance and contextual awareness are balanced in ways that students experience as useful, supportive, and worth continuing to use.","authors":["Madeleine Eastwood","Harshith Narne","Joseph Hilby","Paul Denny","Ashish Aggarwal","Amanpreet Kapoor"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29995","pdf_url":"https://arxiv.org/pdf/2609.29995","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教学助手","编程教育","随机对照试验"],"reason":"研究AI教学助手对学生学习的影响，属于教育技术评估，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.30058","version":1,"title":"Can Labor Markets Function in the Age of AI? The Evaluation Bottleneck in Hiring","zh_title":"AI时代劳动力市场能否正常运转？招聘中的评估瓶颈","abstract":"AI-assisted job-search tools have become increasingly popular by making it easier to find and apply to jobs. But by making it easier for applicants to generate and tailor application materials, they can also reduce how informative those materials are about applicant fit. We study this tradeoff in a hiring market where applicants differ in experience and latent match quality and firms use noisy application materials to decide whom to screen. We ask how AI affects downstream screening and hiring, and which applicants are most adversely affected. As application materials become less informative, a Bayesian firm rationally relies more heavily on coarse observables such as prior experience. Among the four applicant types defined by experience and compatibility for the job, inexperienced-compatible applicants are the most exposed: they lack observable experience and lose the individualized information that could distinguish them from other inexperienced candidates. When screening is costly, these changes can also generate inefficient screening failures in which firms screen no applicants or screen only experienced applicants. We then show that multistage hiring can arise as an endogenous firm response: a relatively inexpensive intermediate assessment allows firms to acquire new evidence of fit before costly full screening. This can restore screening opportunities that disappear under one-stage hiring and give inexperienced-compatible applicants a path to screening. Our results show how AI can shift the central friction in hiring from submitting applications to obtaining credible evaluation, creating entry barriers for high-fit workers without prior experience. Multistage hiring can endogenously arise in response, restoring evaluation opportunities that would otherwise disappear and helping preserve market functioning.","authors":["Itai Ashlagi","Ramesh Johari","Jon Kleinberg","Anushka Murthy"],"categories":["cs.GT","cs.AI","cs.CY","econ.TH"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30058","pdf_url":"https://arxiv.org/pdf/2609.30058","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI招聘","劳动力市场","博弈论"],"reason":"研究AI对招聘市场的影响，无LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.28801","version":1,"title":"The Interface Is Downstream: Designing the Terms of Human-Agent Collaboration","zh_title":"界面在下游：设计人机协作的条款","abstract":"Before an agent responds or acts, much of the experience has already been designed. Memory and retrieval shape what it notices. Evidence rules shape what it may claim. Permissions shape what it can do. Learning rules shape what it carries into the next encounter. The argument comes from Alicia, a personal agent I've built and used since January 2026. A fine-tuning pilot produced no defensible model-performance result. It exposed a provenance failure: Alicia repeated an interpretation from a retrieved synthesis, cited a source note credited by that synthesis, and left the synthesis out of the visible chain. In a model-blind review of thirty-three citations, one reviewer judged that the retrieved intermediary supplied the claim in sixteen relayed citations and part of it in four. Five of twelve citations to directly retrieved targets lacked support in the target excerpt supplied for review. These judgments remain unadjudicated, and the packet is not public. I call the shared setting a humorphic environment: a persistent computational setting that translates a human practice into software. The first Humorphism paper translated partnership. This paper translates the studio, the room where practice happens. The failure prompted an audit of attention, evidence, action, and learning, with a review artifact and available recourse for each. The test is whether the person can inspect and contest what shaped the teammate's behavior. The output is downstream. Correction, consent, and learning carry the collaborative interface back upstream.","authors":["Hector Ouilhet Olmos"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28801","pdf_url":"https://arxiv.org/pdf/2609.28801","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机协作","界面设计","个人代理"],"reason":"论文讨论人机协作界面设计，非用LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2609.28886","version":1,"title":"Characterizing LLM-Based Family Education through the Lens of Activity Theory: A Scoping Review of the HCI Literature","zh_title":"从活动理论视角刻画基于大语言模型的家庭教育：HCI文献的范围综述","abstract":"Large language models (LLMs) are increasingly involved in family education, yet HCI has not systematically explained the educational interactions that emerge around them. This scoping review analyzes 53 HCI studies from 6,540 records across 19 venues. Using activity theory and AODM, it relates participants and educational objects to mediation, labour, and rules. We find that the literature centers on child--parent interaction and on language, AI literacy, and relational learning. The introduction of LLMs enabled conversational, embodied, and spatial systems to generate support from the context of an unfolding interaction. LLMs redistributed educational labour, while family and institutional rules left parents and professionals responsible for interpreting outputs and deciding how they entered practice. Evidence across families and educational purposes remains limited, especially on sustained personalization, repair labour, and how families negotiate authority and rules. The review offers a framework explaining how LLM capabilities become organized through family participation.","authors":["Lan Luo","Yuqi Liang","Jie Cai","Anqi Wang","Dongyijie Pan","Muzhi Zhou","Chun Yu","Pan Hui"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28886","pdf_url":"https://arxiv.org/pdf/2609.28886","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["家庭教育","人机交互","文献综述"],"reason":"该论文是HCI文献综述，关注LLM在家庭教育中的交互设计，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.28990","version":1,"title":"How People Use ChatGPT in Australia: A WildChat Analysis","zh_title":"澳大利亚人如何使用ChatGPT：一项WildChat分析","abstract":"Generative AI chatbots are increasingly embedded in everyday life, yet most large-scale studies describe global patterns. This paper presents an Australia-focused analysis of WildChat, a public dataset of real-world ChatGPT interaction logs. Using descriptive analysis and a multi-layer classification scheme, we analysed 37,845 conversations identified as Australian, examining language diversity, work relevance, interaction intent, topic distribution, turn-taking, temporal change, work activities, and Australia-related domains. Our findings show that the Australian subset is strongly action-oriented and comparatively work-oriented, with most interactions classified as doing and a majority of conversations classified as work-related. The dataset also shows multilingual use and a growing presence of self-expression over time. Australia-related conversations frequently invoke local institutions, laws, regulators, education systems, companies, cultural references, and public services. Finally, we outline implications for future research, including local AI evaluation, multilingual participation, context-aware design, and safeguards for everyday high-stakes domains.","authors":["Ying Ma","Katy Gero","Cl\\'ement Canonne","Craig Jin","Kanchana Thilakarathna"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28990","pdf_url":"https://arxiv.org/pdf/2609.28990","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","使用模式","日志分析"],"reason":"分析真实用户与ChatGPT的对话日志，属于人机交互使用模式研究，不涉及用LL…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.29689","version":1,"title":"Mapping the Authorized Boundary: A Comparative Policy-Vignette Study of Generative AI Governance in Australian Higher Education","zh_title":"划定授权边界：澳大利亚高等教育中生成式AI治理的比较政策情境研究","abstract":"Australian universities regulate students' use of generative artificial intelligence (GenAI) through overlapping policies, procedures, guidance, and assessment instructions, but these environments may classify identical conduct differently. We applied 15 standardized student-use vignettes to the public policy environments of 20 Australian universities and produced 300 university-case classifications. We analyzed binding instruments (Layer A) separately from the full official environment (Layer B), then classified each combination as clearly permitted, permitted with conditions, potential policy breach, clearly prohibited, or indeterminate. We measured cross-university divergence with normalized Shannon entropy. We classified 120 combinations (40.0%) as clearly prohibited, 97 (32.3%) as potential policy breaches, 27 (9.0%) as permitted with conditions, and 56 (18.7%) as indeterminate; none met the strict threshold for clearly permitted. Disclosed language rewriting and a disclosed AI-drafted paragraph produced the highest divergence, whereas an explicit assessment prohibition produced unanimity. Binding instruments remained silent on GenAI in 100 combinations, and guidance clarified 88. The primary researcher led the coding with AI assistance and manually reviewed all 201 queued rows. An independent human second coder, who used no AI assistance, coded a blind 75-row sample. Overall agreement reached 57.3% (unweighted Cohen's kappa = 0.395; bootstrap 95% CI [0.244, 0.539]). The findings distinguish permission-gated from disclosure-based architectures and show that written policies regulate the retention of AI-generated text more clearly than process-only assistance does. Comparative policy-vignette testing evaluates whether written governance supports defensible classifications; it does not predict misconduct or enforcement decisions.","authors":["Biranchi Poudyal"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29689","pdf_url":"https://arxiv.org/pdf/2609.29689","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI治理","高等教育政策","政策分析"],"reason":"研究大学政策对GenAI使用的分类，不涉及用LLM仿真人类被试，而是分析政策文…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.29819","version":1,"title":"Fair Feed Ranking for Participatory Budgeting","zh_title":"参与式预算的公平信息流排序","abstract":"In large-scale participatory budgeting, citizens cannot inspect the full proposal pool, so the order in which proposals are shown becomes a form of agenda-setting power. We argue that fair exposure should therefore be treated as a democratic-design goal. We study Consul Democracy, a widely deployed open-source digital-democracy platform, and show that its proposal feeds are typically ordered by popularity, recency, or comment activity. Building on this diagnosis, we propose FairFeed, a feed-ranking design for PB that uses transparently declared preferences, boosts under-exposed proposals, and admits a rate-limited reject channel for crowd-sourced vetting. We evaluate the design in a simulation anchored in Munich's 2025 PB process and compare it with random, newest, and most-commented feeds. In this simulation, FairFeed broadens proposal discovery, distributes visibility more evenly across the eligible pool, increases cross-cutting support, and improves resistance to manipulation relative to comment-based ranking. We conclude by outlining the human-subjects evaluation needed to test whether onboarding can recover voter preferences accurately enough for deployment in practice.","authors":["Carina I. Hausladen"],"categories":["cs.CY","cs.IR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29819","pdf_url":"https://arxiv.org/pdf/2609.29819","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["参与式预算","排序算法","民主设计"],"reason":"论文研究参与式预算的排序算法，仿真用于评估算法性能，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.28908","version":1,"title":"Automatic Harness Evolution for Hardware Design Verification: Can LLMs Consolidate Gains Across Discovered Harnesses?","zh_title":"硬件设计验证中的自动Harness演化：LLM能否巩固跨发现Harness的收益？","abstract":"Agent behavior depends on the harness surrounding a language model, but it remains unclear whether language models can reliably improve such harnesses for hardware-design tasks. We study automatic harness evolution around a fixed subject model on 12 proprietary design-verification root-cause localization tasks. Across five trials per task, automatically evolved harnesses increased completed attempts by 71-76% and any-hit task coverage by 80-100%, while total correct attempts improved by only 18-24%. The strongest success reproducible at least twice result improved by one task, and later candidates exchanged gains across tasks rather than preserving them. An auxiliary candidate improved on a four-task validation set excluded from search but tied its baseline on a subsequent 12-task replay containing both search and validation tasks, so the selected gain did not persist across the full pool. Across the tested lineage, useful search, evidence, and finalization behaviors appeared in different candidates but did not consistently consolidate into a single harness that dominated across tasks and metrics. In a separate CVDP cross-benchmark case study, an automatically evolved defined-width repair harness produced 35.6% more functional passes than its 142-task reference baseline; the final functional verifier scored completed outputs but was not shown to the subject agent during repair. These results support archive-aware selection when evolution yields complementary specializations without consistent consolidation.","authors":["Kidus Seyoum","Ajay Mittur"],"categories":["cs.SE","cs.LG"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28908","pdf_url":"https://arxiv.org/pdf/2609.28908","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["硬件验证","LLM agent","自动演化"],"reason":"研究LLM自动改进硬件验证harness，属多智能体协作解题，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29016","version":1,"title":"EvoTreeNAD: Genealogy-Guided Evolution for LLM-Driven Neural Architecture Discovery","zh_title":"EvoTreeNAD：谱系引导的进化算法用于LLM驱动的神经架构发现","abstract":"AI-driven scientific discovery accelerates research by autonomously developing solutions and designs. Large language model (LLM) agents support this process through iterative generation and evaluation. Yet these iterations alone do not ensure cumulative progress or establish which directions to pursue next. Costly evaluation further constrains the scope of exploration. Neural architecture discovery brings these challenges together, coupling open-ended design with resource-intensive experimentation. We introduce EvoTreeNAD, a genealogy-guided evolutionary algorithm that constructs trainable architectures without a supplied seed or a hand-specified search space. Starting from an empty root, it grows a persistent genealogy in which each new node represents a complete architecture. Top-percentile values computed from each node and its descendants guide lineage selection. Using the selected design history, an Idea Agent proposes a variant and a Code Agent implements it. Each evaluated variant becomes a child node, expanding the genealogy while providing evidence for subsequent lineage selection. Our theoretical analysis establishes the existence of stationary variation regimes as the genealogy grows. Under specified variation assumptions, sustained top-percentile family values quantify the probability of generating high-reward architectures in these regimes. EvoTreeNAD discovers architectures that outperform the compared NAS and NAD baselines, achieving CIFAR-10/100 test errors of $2.05{\\pm}0.06\\%$ and $15.09{\\pm}0.22\\%$. On all six MedMNIST-v2 tasks, the discovered architectures surpass the strongest listed baselines. A controlled CIFAR-10 study further shows that EvoTreeNAD outperforms direct generation, best-of-$N$ greedy continuation, and full-family-mean routing.","authors":["Lishan Yu","Derek Jiu","Qizhen Lan","Xiaoqian Jiang"],"categories":["cs.NE","cs.LG"],"primary_category":"cs.NE","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29016","pdf_url":"https://arxiv.org/pdf/2609.29016","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["神经架构搜索","多智能体系统","进化算法"],"reason":"多智能体协作解决神经架构搜索，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.27535","version":1,"title":"KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration","zh_title":"KITE：通过稀疏旗舰校准扩展Jev人口实验","abstract":"KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.","authors":["Hengyu Li (The University of Tokyo)"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27535","pdf_url":"https://arxiv.org/pdf/2609.27535","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类数据对照","政策评估"],"reason":"用LLM仿真人类被试，有真实人类数据对照，评估偏差并传播不确定性，用于政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":5,"question":"如何在大规模人口实验中，用廉价的类型化行为核模型结合稀疏的旗舰模型校准，来估计干预效应并传播人类-模型差异的不确定性？","design":"KITE 使用 TypeSafe 的 Jev 行为核模型（类型化行为核）对每个唯一状态查询一次，生成决策分布，然后用表格化执行模拟任意规模的人口，并通过事件键控随机数和共同随机数控制变异性；昂贵的旗舰模型仅用于稀疏的配对锚点，以估计干预效应；人类-模型差异被作为共享误差传播到所有结论中。","baseline":"对照的真实人类数据包括 Epstein 实验（9,070 名参与者）和 37 个留出的 SocSci210 实验，以及一个 16 国研究中的内容保真度标准。","findings":"在 Epstein 实验中，覆盖 1.7% 状态的锚点将效应误差降低了 41%（绝对 MAE 降低 0.0125）；在 37 个留出的 SocSci210 实验中，0.5-1.5% 的锚点覆盖率将捕获的决策增益从 0.27 提高到 0.39。共享差异传播在名义 80% 和 90% 的置信水平下分别实现了 93% 和 96% 的回顾性覆盖率，而仅使用人类抽样不确定性时分别为 29% 和 36%。","reliability":"论文承认的局限包括：仅有两个留出的人类参考数据集测试混合方法；效应幅度需要进一步校准，差异斜率在不同研究选择间变化；稀疏校准可能引入旗舰模型误差；核模型可能过度预测规范信息；实验不授权预测文献中不存在的干预；记忆一致性是局部的，队列级边际校正可能抹去真实的持续性或处理路径；刺激重建、省略卡片图像、人口统计压缩等限制了保真度；专有模型和访问限制阻碍了可重复性。","relevance":"该研究直接针对用 LLM 仿真人类被试的核心问题，提供了与真实人类数据对照的验证，并传播不确定性，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"该方法通过稀疏旗舰模型校准和共享误差传播，在保持低成本的同时提高了效应估计的准确性，值得借鉴其校准策略和不确定性量化方法。｜可迁移到政策评估场景，如税收政策变化对劳动供给的影响、福利项目对消费行为的影响，或信息干预对金融决策的影响。｜设计一个实验：用 Jev 核模型模拟不同人口群体对政策公告的反应，以真实调查数据（如消费者预期调查）为基准，施加政策处理（如利率变化信息），测量预期调整和消费意愿，并用稀疏旗舰模型校准关键状态，传播人类-模型差异。"}},{"id":"2609.28470","version":1,"title":"StudentBench: AI and human tutoring yield equivalent GRE learning gains","zh_title":"StudentBench：AI与人类辅导在GRE学习收益上等效","abstract":"Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.","authors":["Curtis Northcutt","Inaara Hasmani","Kevin Feng","Trevor Khangi","Andreas Plesner","Jonas Mueller"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28470","pdf_url":"https://arxiv.org/pdf/2609.28470","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育实验","人类对照"],"reason":"用LLM替代人类导师进行教学实验，并与人类导师对照，评估学习效果，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":6,"question":"大语言模型（LLM）作为AI导师能否在GRE学习收益上达到与人类专家导师统计等效的效果？","design":"用多个LLM（如Gemma 4 31B等）扮演AI导师，对2383名人类参与者进行GRE定量和语文部分的辅导，同时设置人类导师辅导组和无辅导对照组，测量学习收益（前后测成绩提升百分比），并收集175,000条学生-AI消息分析对话行为。","baseline":"人类专家导师辅导组的学习收益数据，以及无辅导对照组的学习收益数据。","findings":"AI辅导在GRE学习收益上与人类专家辅导统计等效（p=.015），且在七个GRE领域中五个领域的最佳AI导师平均超过人类导师。一个AI导师（Gemma 4 31B）以低918倍的成本实现了与人类辅导等效的学习收益（p=.044）。","reliability":"论文承认局限：只测量即时学习收益，未评估长期保持；参与者均为能读写英语的成年人，未测试跨语言、设备或教育环境；未设置学生独自练习的对照组，无法分离AI交互的额外收益；排行榜评估基于专家评价和对话行为，而非实际学习效果。","relevance":"该研究用LLM替代人类导师进行教学实验，并与真实人类导师及无辅导组对照，评估学习效果，属于典型的人类仿真研究，且提供了大规模真实人类数据作为基准，值得精读以了解仿真等效性检验的设计与局限。","inspiration":"借鉴其多组对照设计（AI处理、人类处理、无处理）和统计等效性检验（TOST）来严格评估AI干预是否非劣于人类专家。｜可迁移到金融教育或投资者决策辅导场景，例如测试AI投教助手能否在提升投资者金融素养或改善投资决策上达到人类理财顾问的效果。｜以真实投资者为被试，随机分为AI投教组、人类顾问组和无辅导组，处理为一段时间的个性化金融知识辅导，结果变量为金融素养测试得分或模拟投资组合表现，对照真实人类顾问组和无辅导组的数据，并采用等效性检验。"}},{"id":"2609.28372","version":1,"title":"Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer","zh_title":"算法购物：代理式AI如何将人类启发式用作替代消费者","abstract":"Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using \"Tool-Lab,\" an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pricing and promotional framing) influence AI shopping agents. Across eight commercially deployed LLMs from three providers, we trace pre-choice information acquisition. Under zero cost, pricing cues rarely mislead. Imposing acquisition costs under a vague goal prompt leads LLMs to omit diagnostic attributes required to compute unit price and choose suboptimal choices resembling human heuristics. Relative to a specific goal prompt that mainly preserves diagnostic search and choice optimality, a vague goal prompt under constraints creates a search-mediated vulnerability. This research demonstrates that marketing heuristics in delegated AI shopping are governed by storefront information architecture, not necessarily immutable LLM flaws.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28372","pdf_url":"https://arxiv.org/pdf/2609.28372","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为代理消费者模拟人类购物决策，并与人类启发式对照，涉及营销实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":9,"question":"在委托AI代理购物时，营销定价线索（如尾数定价和促销框架）如何通过信息获取成本与目标提示的具体性影响LLM的信息搜索和选择最优性？","design":"使用Tool-Lab实验范式，将产品属性隐藏在需要付费的工具调用之后，操纵信息获取成本（0、1、5美分）和提示目标具体性（模糊：“找最划算的” vs. 具体：“找每盎司最低价”），对8个商用LLM（来自Google、OpenAI、Anthropic）进行咖啡选择实验，测量信息搜索深度、搜索组成和选择最优性。","baseline":"无对照","findings":"在零成本或具体目标提示下，LLM大多做出规范最优选择；但在获取成本存在且目标模糊时，多数LLM会减少搜索深度，省略诊断性属性（如美分或重量），导致次优选择，类似于人类启发式决策。","reliability":"论文未讨论","relevance":"本研究通过实验操纵环境约束（成本与提示），揭示了LLM启发式行为的条件性，为评估LLM作为人类被试替代品的可靠性提供了关键证据，值得精读原文以理解其方法细节和边界条件。","inspiration":"借鉴其通过工具调用施加信息获取成本并操纵提示具体性的设计，可迁移到消费者金融决策（如信用卡选择、贷款比较）或投资者信息处理场景；例如，用LLM模拟投资者在获取公司财务指标需付出成本时，模糊目标（“选只好股票”）与具体目标（“选市盈率最低的股票”）下的信息搜索与选择，并与真实投资者眼动或点击流数据对照。"}},{"id":"2608.22859","version":2,"title":"WARP: Wasserstein-Aligned RAG for Population Opinions","zh_title":"WARP：面向群体意见的Wasserstein对齐检索增强生成","abstract":"RAG systems are increasingly used to summarize what large collections of documents say. A user asks \"What do people think about X?\" and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread retrieved documents apart, but with no target distribution to aim for. Calibration methods based on KL or JS divergence do target one, yet treat opinion bins as unordered: confusing strong positive with strong negative costs no more than an adjacent-bin miss. We introduce WARP, a family of post-retrieval algorithms that calibrate retrieved evidence to the population's opinion distribution. WARP first recovers underrepresented opinions that cosine ranking may bury, then uses Wasserstein-1 distance to select documents whose sentiment-intensity distribution matches the population target, capturing the ordinal structure ignored by KL and JS divergence. We develop three variants for dense, sparse, and variable candidate pools, trading off calibration quality and speed. Across three review domains spanning 35K documents, 156 queries, and 26 entities, WARP's domain-matched variants reduce distributional error by at least 43% with sub-second latency. These gains carry through to generation: a five-judge LLM panel prefers WARP-generated answers in 86% of decided comparisons at k <= 5.","authors":["Aman Singh Thakur","Aditya Agrawal","Alwarappan Nakkiran","Alex Karlsson"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-24","first_seen":"2026-08-25","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.22859","pdf_url":"https://arxiv.org/pdf/2608.22859","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","意见分布校准","RAG"],"reason":"用LLM生成代表人群意见的摘要，并与真实意见分布对齐，属于仿真人类态度，且有真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":10,"question":"如何让检索增强生成（RAG）系统在回答关于人群意见的查询时，所选证据的意见分布忠实于总体人群的意见分布，避免多数意见淹没少数意见？","design":"该论文提出 WARP 算法族，用于在 RAG 流程的后检索阶段校准所选文档的意见分布。它首先通过缺陷感知的池扩展策略恢复被余弦相似度排序埋没的少数意见文档，然后使用 Wasserstein-1 距离作为度量，贪婪地选择文档子集，使其情感强度分布与目标人群分布匹配。论文在三个评论数据集（Amazon 卖家论坛、Yelp 酒店评论、OpinRank 汽车评论）上进行了实验，共 35K 文档、156 个查询、26 个实体，比较了 WARP 与 Top-k、MMR、DPP、OpinionMMR、KL/JS 校准等基线在分布误差、实体匹配率和延迟上的表现，并通过 LLM 评委小组评估生成答案的质量。","baseline":"论文使用从语料库中提取的实体级情感强度分布作为目标人群分布，该分布由 LLM 对每个文档进行情感强度标注后聚合得到，作为真实人群意见分布的代理。","findings":"WARP 的领域匹配变体在三个评论领域中将分布误差降低了至少 43%，且重排序延迟低于 310 毫秒（p99）。在生成评估中，五位 LLM 评委在 k≤5 时对 WARP 生成答案的偏好率高达 86%。","reliability":"论文指出，WARP 的性能依赖于实体密度：在实体稀疏的领域，需要混合变体（如 λ-MMR 或 WassRank OT）来平衡相关性与校准；此外，目标分布是从语料库中估计的，可能无法完全代表真实人群，但论文未深入讨论这一局限。","relevance":"该论文直接针对用 LLM 生成代表人群意见的摘要问题，通过 Wasserstein 距离对齐意见分布，属于仿真人类态度分布的研究，且有真实评论数据作为基准，值得阅读原文以了解其算法细节和评估方法。","inspiration":"该方法借鉴了将检索证据校准到目标分布的思想，并利用 Wasserstein 距离捕捉有序情感结构，可用于经济金融领域中需要从文本中提取群体意见或情绪分布的场景。｜可以迁移到消费者信心指数构建、财报电话会议情绪分析、社交媒体上的政策预期形成等场景，其中需要从大量文本中估计人群的态度分布。｜一个可行的研究设计是：以 Twitter 上关于某经济政策（如加息）的推文为语料，用 LLM 对每条推文进行情感强度标注（如从强烈反对到强烈支持），得到目标分布；然后使用 WARP 算法从检索到的推文中选择一小部分作为 LLM 生成摘要的证据，确保所选推文的情感分布与总体分布一致；最后将生成的摘要与基于随机抽样或简单检索的摘要进行对比，评估其在预测真实调查数据（如密歇根消费者信心指数）上的准确性。"}},{"id":"2609.27165","version":1,"title":"Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement","zh_title":"计数证据而非句子：面向长文本价值测量的LLM判断调和证据融合","abstract":"Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.","authors":["Yuhe Wu","Rui Qian","Guangyu Wang","Yuran Chen","Yuanchao Zhu","Junjie Yang","Zhengheng Li","Jiulin Cai","Tianyi Zhang","Zihan Dong","Jiaxin Liu","Yujie Chen","Guang Zhang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27165","pdf_url":"https://arxiv.org/pdf/2609.27165","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM价值测量","证据融合","长文本分析"],"reason":"用LLM测量公众价值取向，有真实人类标注数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":13,"question":"如何融合长文本中句子级LLM判断，使决定性证据主导最终价值取向测量，同时抑制不确定句子？","design":"本研究不是人类仿真实验，而是提出一种训练无关的决策融合规则TEF，用于融合句子级LLM判断来测量长社交媒体帖子的价值取向。具体做法：将长文本分割为句子，用LLM对每个句子输出两个立场标签的对数概率，计算对数几率，并用归一化信息增益加权，再求和得到文档级得分。在MIND基准（8,358条中英文帖子，六个价值维度）上，用五个LLM（Qwen2.5-7B、LLaMA3-8B、Qwen3-14B、DeepSeek-V3.2、GPT-4o-mini）评估TEF与直接预测、多数投票、软投票的性能。","baseline":"有真实人类标注数据：MIND基准包含8,358条中英文社交媒体帖子，由人类标注了六个价值维度上的立场标签。","findings":"TEF在五个LLM和两种语言上平均比最强基线（直接预测、多数投票、软投票）高出4.5个准确率点和4.6个宏F1点。TEF在Qwen2.5-7B上校准误差最低，且对提示词扰动更稳健。","reliability":"论文未明确讨论失效条件，但指出直接求和LLM logits不可靠，因为next-token概率只是近似校准，尤其在最大不确定点附近；消融实验显示，在对数几率变换和熵加权同时移除时，某些模型上性能会低于软投票，表明两者互补。","relevance":"该研究虽非人类仿真实验，但提供了用LLM测量公众价值取向的可靠方法，且有真实人类标注数据对照，可作为仿真研究中测量态度或价值观的工具，值得阅读原文了解融合规则细节。","inspiration":"借鉴其决策融合思路，将长文本分解为句子级判断并用信息增益加权融合，可提高LLM对复杂文本的测量准确性｜可迁移到经济金融领域的文本测量，如从财报电话会议记录中提取管理层情绪、从新闻中测量政策不确定性、从社交媒体帖子中测量消费者信心或通胀预期｜设计：用LLM对财报电话会议记录的每个句子判断管理层语气（积极/消极），用TEF融合句子级判断得到文档级情绪得分，以分析师一致预期或后续股票收益作为真实数据对照，评估LLM情绪测量的预测效度。"}},{"id":"2609.26861","version":1,"title":"Rule-Based Pricing Algorithms and Market Outcomes: An Experimental Study","zh_title":"基于规则的定价算法与市场结果：一项实验研究","abstract":"Rule-based pricing tools are widespread in digital commerce, yet we know little about how their design shapes market outcomes. In a controlled market experiment, participants use dashboards to build pricing algorithms competing in a sequential Bertrand game over multiple periods. We vary design features commonly found in commercial repricing tools: warnings about price wars, pre-configured strategies, and advice from a large language model. Most treatment variations raise market prices with effects driven by an increase in starting prices and more cooperative algorithm designs. The results matter for competition policy, platform regulation and current discussions on regulating algorithm design tools.","authors":["Adrian Hillenbrand","Hans-Theo Normann","Matthias Potarca","Tobias Werner"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26861","pdf_url":"https://arxiv.org/pdf/2609.26861","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM建议","经济实验","算法定价"],"reason":"用LLM提供建议影响人类定价实验，有真实人类数据对照，属经济实验场景，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":12,"question":"规则型定价算法的设计特征（价格战警告、预配置策略、LLM建议）如何影响市场结果与合谋行为？","design":"受控市场实验：人类被试通过仪表盘构建定价算法，在序贯伯特兰博弈中竞争50期，共5轮超级博弈；处理为三种设计特征：盈利性提示（警告价格战）、LLM建议、预配置策略菜单与默认合作性价格匹配；结果变量为市场价格、起始价格、算法合作性设计。","baseline":"基线处理（无设计特征）作为对照，比较不同处理组与基线组的价格和算法设计差异。","findings":"大多数处理变体提高了市场价格，效应主要由起始价格提高和更合作的算法设计驱动。LLM建议和预配置策略等设计特征可能促进合谋，对竞争政策和平台监管有启示。","reliability":"论文未讨论","relevance":"该研究用LLM提供建议影响人类定价实验，有真实人类数据对照，属经济实验场景，可迁移到LLM仿真人类决策的可靠性评估，值得读原文了解LLM建议的具体效果。","inspiration":"借鉴其将LLM建议作为处理变量嵌入人类实验、并观察对策略选择和均衡结果影响的设计｜可迁移到资产定价实验或政策公告预期形成场景，研究LLM建议对投资者行为或公众预期的影响｜以人类被试为对象，处理为是否提供LLM投资建议，结果变量为报价或预期值，对照真实市场数据或调查数据评估LLM建议的偏差与可靠性"}},{"id":"2609.27639","version":1,"title":"Agent-based Modeling: Equilibrium, Echo Chambers, and Efficiency in Hybrid Coevolutionary Opinion Games","zh_title":"基于智能体的建模：混合协同演化观点博弈中的均衡、回音室与效率","abstract":"Opinion formation in online networks involves changes in both beliefs and social ties. Analytical models make it possible to study equilibrium and social cost, but usually represent communication as a fixed numerical update. LLM-driven agents offer a language-based alternative, yet their convergence and collective efficiency remain unclear. We develop the Hybrid Coevolutionary Opinion Game (H-COG), combining cost-minimizing Friedkin-Johnsen agents (Type-C) and Phi-4 language agents (Type-L) in a dynamically rewired K-nearest-neighbor network. We initialize 50 agents with opinions drawn from 5,199 Reddit comments on gun control and abortion. The comments are scored on a continuous [-1,+1] scale using a fine-tuned RoBERTa regressor, and a mixing parameter sets the proportion of each agent type. The experiments cover nine population compositions, three initial network topologies, and two topics. All 540 runs meet the convergence criterion within the simulation horizon. Under Type-L updating, the coevolving network reaches an attractor as reliably as it does under the analytical update rule, making an equilibrium-based efficiency comparison possible. The pooled Price of Anarchy is $5.558 \\pm 0.309$ for purely Type-L populations, compared with $1.139 \\pm 0.005$ for purely Type-C populations. A decomposition of social cost attributes most of this gap to language agents moving away from their intrinsic opinions, rather than to greater disagreement with their neighbors. The main findings are consistent across the three initial network topologies.","authors":["Ming-Zhi Jiang","An-Tzi Teng","Jun-En Liu","Po-An Chen","Yung-Ming Li"],"categories":["cs.GT","cs.MA","cs.SI"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27639","pdf_url":"https://arxiv.org/pdf/2609.27639","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","观点动力学","社会模拟"],"reason":"用LLM agent模拟观点动态，但无真实人类行为对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.27686","version":1,"title":"Mining Meaning: Measurement Error in AI-Assisted Literature Reviews","zh_title":"挖掘意义：AI辅助文献综述中的测量误差","abstract":"Researchers increasingly use generative AI, particularly large language models (LLMs), to automate tasks across the research pipeline. We study the reliability of these tools at the reading, classification, and synthesis of large bodies of academic literature. We frame LLM-assisted literature reviews as a measurement problem, treating models as measurement systems and tracing how their errors affect downstream conclusions. As a test case, we use three different implementations of ChatGPT to identify and extract metadata from economics papers that use rainfall as an instrumental variable. We benchmark each implementation against a subset of human-labeled evaluation data, and then deploy those implementations to extract metadata from the full corpus. The LLMs perform well on binary classification, but performance deteriorates as tasks demand greater contextual interpretation. More importantly, how much researchers can rely on model outputs depends not only on the complexity of the reading task but also on the type of claims the data is asked to support. The same amount of measurement error substantially affects paper-level claims while having little effect on broader claims about the literature. Measurement error in LLM-generated data is thus most consequential at precisely the level of detail that constitutes an LLM's principal value added over human reviewers. We conclude that standard model performance metrics are informative about the quality of generated data but do not by themselves establish the credibility of downstream inference. Researchers must also evaluate whether substantive claims are robust to the measurement system used to generate the underlying data.","authors":["Jeffrey D. Michler","Kieran Douglas","Anna Josephson"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27686","pdf_url":"https://arxiv.org/pdf/2609.27686","source_feed":"econ.GN","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","测量误差","文献综述"],"reason":"LLM替代人工标注文献，属标注员替代而非仿真被试，但涉及测量误差与可靠性，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2608.06968","version":2,"title":"How a shared state is described determines whether AI agents synchronize","zh_title":"共享状态的描述方式决定AI智能体是否同步","abstract":"Language-model agents increasingly act in populations, where the outcome that matters is collective: whether they align, split or fail to coordinate. Each acts not on the world but on a text description of it, a choice usually fixed in software. Using synchronization, the canonical probe of how interaction rules produce collective order, we show that this choice can decide the outcome. Agents on a circle chose to advance, stay or move back after reading the others' relative positions, in 507,112 valid responses across matched populations, controlled inputs and three model families. In GPT, numerical summaries aligned every matched population at both positive couplings, whereas histograms aligned none; Claude showed the reverse at the stronger coupling. Re-describing identical states shifted action probabilities in all three families, even between histograms carrying the same information. No single directional coefficient explained the outcome: state descriptions are part of the interaction rule that turns individual responses into collective order.","authors":["Takahiro Ezaki","Naoto Imura","Katsuhiro Nishinari"],"categories":["physics.soc-ph","cs.AI","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-09-24","first_seen":"2026-08-10","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.06968","pdf_url":"https://arxiv.org/pdf/2608.06968","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","集体行为","同步"],"reason":"LLM agent群体同步实验，无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.24574","version":2,"title":"Evaluating Decision Models for Text Annotation in Computational Social Science","zh_title":"评估计算社会科学中文本标注的决策模型","abstract":"Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the labels generated by such models. Decision models, a new model class built for categorical question answering, answer typed questions with a choice, a probability distribution over the label set, and a confidence score rather than free text, at a small fraction of frontier inference prices. Whether their answers are accurate, and whether that stated confidence can be trusted on social science constructs, are unknown. Here, we mirror the evaluation of Ziems et al. (2024) on 18 computational social science classification tasks (7,977 items), comparing the first commercial decision model and two open-weight counterparts against 19 frontier and open-weight language models under the same zero-shot protocol, and extending the decision-model comparison to eleven open-weight systems released in the week after it. The decision model trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost. Its confidence is better calibrated than the verbalized confidence of 16 of the 19 LLMs, yet three frontier models show lower median calibration error (0.157 against 0.066). While items above 0.9 confidence are typically labeled accurately (median accuracy 0.815), on one task, empathy in peer-support dialogues, the model reports high confidence while performing near chance. Nonetheless, our results suggest that decision models are useful as a first step in the annotation pipeline: routing low-confidence items to an LLM matches or exceeds the LLM alone at a quarter to half of its cost.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-24","first_seen":"2026-09-22","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2609.24574","pdf_url":"https://arxiv.org/pdf/2609.24574","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","计算社会科学","模型评估"],"reason":"评估LLM用于文本标注，替代人工标注员，属于D1边界情形，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.26926","version":1,"title":"Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation","zh_title":"专家在LLM分歧处崛起：利用跨模型分歧定位专家精力以修订大规模标注的LLM编码手册","abstract":"Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by cross-LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring-session transcripts show that Rationale Labeling yielded the highest LLM-labeling accuracy (64.9%) against expert labels, outperforming the expert-revised codebook (57.8%). The best Question Answering setting also outperformed it (60.5%). Our work shows that LLMs can be used to strategically target expert attention, shortening months of codebook revision to days without sacrificing labeling performance.","authors":["Zeyu He","Zhuqian Zhou","Kirk Vanacore","Rene F. Kizilcec","Ting-Hao 'Kenneth' Huang"],"categories":["cs.CL","cs.AI","cs.HC","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26926","pdf_url":"https://arxiv.org/pdf/2609.26926","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","编码手册修订","人机协作"],"reason":"用LLM辅助标注，非仿真人类被试，但涉及专家反馈与LLM协作，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27043","version":1,"title":"EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues","zh_title":"EduBehaviors：基于断言的模式用于教育对话的可审计编码","abstract":"Large language models have allowed the rapid deployment of pedagogical annotations corresponding to constructs of interest, allowing a natural language interface for generating classifications on a conversational dataset. However due to the opaque nature of LLM reasoning, we have no verifiable, mechanistic insight into why a model chose a label for an utterance. We introduce the EduBehaviors framework, an interpretable, scalable approach to annotating educational data that uses LLMs to measure repeated observable behaviors relevant to many constructs of interest and then learns a classifier for the construct based on these observable behaviors. We evaluate the framework on the TalkMoves dataset, predicting the Teacher TalkMoves labels. Our best configuration results in a macro-F1 of 0.673 and 0.688 Cohen's kappa, proving competitive with direct prompting approaches. In addition, we release EduBehaviors Toolkit, two tools allowing researchers to operationalize the EduBehaviors framework in their own data.","authors":["Julian Bernado","Ana Trindade Ribeiro","Xander Beberman","Susanna Loeb"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27043","pdf_url":"https://arxiv.org/pdf/2609.27043","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","教育对话","可解释性"],"reason":"用LLM标注教育对话，替代人工标注，非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27811","version":1,"title":"A Decade of Climate Polarization on Brazilian YouTube using Language Models","zh_title":"巴西YouTube上气候极化十年研究：基于语言模型","abstract":"Online platforms have become arenas for the public contestation of climate change, shaping how scientific knowledge, denial, and uncertainty are expressed and disputed. Yet longitudinal evidence remains limited for YouTube, especially for Portuguese-language discourse. Addressing this gap, we characterize how climate stances are expressed and contested over time in a large corpus of Portuguese-language YouTube comments retrieved through Brazil-oriented climate-related searches. To support this analysis in a noisy, imbalanced, and low-resource setting, we collect more than 240,000 comments posted between 2014 and 2024 and formulate stance detection as a three-way classification task (Believer, Denier, and Inconclusive). We operationalize stance attribution through a scalable self-training pipeline based on Llama 3.1, using Low-Rank Adaptation (LoRA) and hybrid instance selection to expand the training set with high-confidence pseudo-labeled examples while preserving class diversity. This approach improves coverage and class balance for minority and rhetorically complex classes, enabling large-scale stance attribution without extensive manual annotation. Our results show that polarisation is marked by interactional asymmetries: denialist comments are less prevalent, but they are associated with a comparatively higher share of cross-stance contestation, while pro-consensus discourse is more strongly reinforced within stance-homogeneous threads.","authors":["Daniel Morais","Diego H. M. Magalhaes","Gabriel H. Silva","Andrea Failla","Valeria de C. Santos","Helen C. S. C. Lima","Carlos H. G. Ferreira"],"categories":["cs.SI","cs.CL","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27811","pdf_url":"https://arxiv.org/pdf/2609.27811","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["立场检测","LLM标注","社交媒体分析"],"reason":"用LLM做立场标注，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.27327","version":1,"title":"Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows","zh_title":"视觉语言模型能否分析以人为中心的视频？映射模型能力与人机协作工作流","abstract":"Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.","authors":["Xiyuan Shen","Jiuyang Lyu","Seokhyun Hwang","Huanfen Yao","Shwetak Patel","Zhihan Zhang","Jacob O. Wobbrock"],"categories":["cs.CV","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27327","pdf_url":"https://arxiv.org/pdf/2609.27327","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["VLM","视频标注","人机协作"],"reason":"用VLM替代人工标注视频，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.27063","version":1,"title":"Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses","zh_title":"数据科学课程中学生使用LLM及AI生成题目难度的局限性","abstract":"This paper presents a multi-source classroom study conducted during a 10-week quarter in data science courses at Drexel University. We first investigate the behaviors of student engagement with large language models (LLMs) using four surveys across three data science courses. Second, we evaluate the construct validity of multiple-choice questions (MCQs) generated by an LLM for in-lecture retrieval practice. Based on 378 authored questions (311 deployed, producing 7{,}888 student responses), we analyze whether the difficulty ratings assigned by an LLM match empirical item difficulty. Our study shows that student engagement with LLMs varied across courses and increased over the term. Although students expressed high satisfaction and reported saving considerable time, their perception of deep learning benefits declined, and many noted a tendency toward over-reliance. Regarding the difficulty ratings of LLM-generated MCQs, the Easy, Medium, and Hard labels correlated closely with its assigned Bloom's Taxonomy levels (Spearman $\\rho=0.90$), reflecting an artifact of co-generation. However, neither metric predicted empirical item difficulty (difficulty label $\\rho=0.06$; Bloom level $\\rho=0.02$). The ratings reflect the structural formatting of a question rather than its underlying difficulty.","authors":["Yuan An","Lei Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27063","pdf_url":"https://arxiv.org/pdf/2609.27063","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM生成题目","教育评估","难度预测"],"reason":"LLM生成题目并评估难度，替代教师出题，非仿真人类被试，但涉及LLM与人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.28114","version":1,"title":"Watching What We Eat: Information Quality and Body Image in Diet-Related YouTube Videos","zh_title":"观看我们所吃的：饮食相关YouTube视频中的信息质量与身体意象","abstract":"The widespread use of social media, particularly image- and video-based platforms, has turned them into key sources of both normative and informational content related to health and diet. This may contribute to the development of disordered eating behaviors or, potentially, eating disorders. This study uses mixed-methods analysis applied to 3129 YouTube videos about diet and weight loss in order to quantify the level of risk of low-quality information and heightened focus on the body image. We adapt three quality measurement frameworks from the literature -- PRHISM, HONcode and SMEC -- to the online video context, perform manual annotation of a sample of the data, and design an LLM content characterization pipeline to score the videos on the quality of their content and the focus on the human body. Surprisingly, we find that videos around personal storytelling and mindset & motivation are associated with higher-quality content, whereas supplement reviews (arguably more medically sensitive ones) are not. Further, the body-related mentions of weight measurement and negative body image are associated with an increased viewership, whereas the mentions of positive body image are associated with an increased engagement rate in terms of likes and comments, but not viewership. Worryingly, we find a cluster of videos categorized as \"music\" which promote the dietary supplements Mitolyn and the injectable weight loss drug Mounjaro. As video-based platforms grow in popularity, particularly among younger audiences, studies such as the one presented here are essential for developing empirically grounded tools to enhance the detection of harmful content and inform more effective moderation practices.","authors":["Maddalena Ghiotti","Daniela Paolotti","Yelena Mejova"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28114","pdf_url":"https://arxiv.org/pdf/2609.28114","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","内容分析","社交媒体"],"reason":"用LLM做内容标注，替代人工编码，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28362","version":1,"title":"Threat Amplified, Blame Restrained: LLM-Assisted Media Framing Analysis of the 2026 Bangladesh Measles Outbreak","zh_title":"威胁放大，指责克制：2026年孟加拉国麻疹疫情的LLM辅助媒体框架分析","abstract":"How news media frame and emotionally code a public health emergency shapes public risk perception and trust, yet outbreak-coverage dynamics remain understudied for low- and middle-income countries (LMICs). We examine sentiment and stance in English-language Bangladeshi coverage of the 2026 measles outbreak -- the country's most severe in two decades, with over 97,000 suspected cases and 600 deaths across 61 of 64 districts, unfolding after the 2024 change of government and a 2024-2025 vaccine stockout. Using the Internet Archive, we build a reproducible corpus of 403 headlines from seven national outlets (396 in-window in 2026), label them for binary sentiment and four-way stance via a large language model under a locked codebook, and validate against a two-coder human-adjudicated gold standard (n=153; Cohen's kappa=0.89 stance, 0.75 sentiment). Aligned to the DGHS epidemic curve, coverage grew significantly more negative (56% to 88% negative; Cochran-Armitage z=4.12, p<.001) and risk-amplification framing intensified (44% to 84%; z=4.15, p<.001). Media negativity lagged incidence, tracking cumulative mortality. Contrary to the political backdrop, blame remained a minority frame (~9% overall) and was overwhelmingly systemic (32 of 37, 86%) rather than directed at named actors. The pipeline offers a scalable, transparent method for LMIC outbreak-media analysis; Bangladeshi coverage amplified threat far more than it assigned political blame.","authors":["Shahan Ahmed"],"categories":["cs.SI","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28362","pdf_url":"https://arxiv.org/pdf/2609.28362","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","媒体框架分析","公共卫生"],"reason":"用LLM做媒体框架标注，替代人工编码，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.28388","version":1,"title":"OranSim: Simulating Social Media Marketing","zh_title":"OranSim：模拟社交媒体营销","abstract":"Social simulation studies how individual behavior and social interaction produce collective outcomes. In social media marketing, campaign actions shape which consumers encounter the content and how they respond; these responses then spread through the population. We propose OranSim, a social simulation framework that connects creative, creator, targeting, and budget choices to this process. Heterogeneous consumers receive exposure according to content matching and platform allocation and generate initial responses, which propagate among 60 population segments. Candidate campaigns share the initial population and aligned random numbers, making their response trajectories comparable under action changes. In a controlled synthetic campaign, doubling the budget approximately doubles reach while lowering mean content match and engagement probability among the reached consumers; mean 14-day cumulative simulated response mass rises to 1.96 times the baseline. LightGBM predictors fitted to 39,000 historical RedNote notes estimate platform engagement with log-scale $R^2$ of 0.56--0.62 in five-fold cross-validation; a separate 12,154-note corpus supplies temporal, unseen-creator, and held-out-niche test splits. Public-data experiments evaluate policy value and audience ranking, and paired synthetic outcomes test counterfactual scoring. Together, scenario trajectories and engagement estimates support campaign selection according to a prespecified marketing objective. Code is available at https://github.com/OranAi-Ltd/oransim.","authors":["Jianxiang Ma","Mingfu Zhang","Xiaocui Yang","Yichen Gao","Junzhao Huang","Yuesong Hou"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28388","pdf_url":"https://arxiv.org/pdf/2609.28388","source_feed":"cs.SI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","营销仿真","无LLM被试"],"reason":"社会模拟但无LLM作为被试，且无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.27994","version":1,"title":"Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance","zh_title":"合规于局部控制，集体歧视：受监管金融中多智能体AI的治理架构","abstract":"Financial institutions are beginning to deploy agentic workflows in credit, fraud, collections, compliance, and operational control. Governance remains largely component-centric: each model or agent is specified, tested, authorized, and monitored locally. That is insufficient when institutional risk arises from the joint behavior of many locally acceptable components. We call this gap constitutional non-compositionality: local compliance checks need not compose into acceptable collective outcomes such as bounded disparate impact, market integrity, or traceable accountability. We propose ARIA as a finance-specific reference architecture and falsifiable research agenda for agent-population governance. It organizes six capabilities across normative-accountability, execution-control, and assurance-learning planes: policy specification, population-level observed-versus-expected behavior monitoring (M2), bounded authority, runtime containment, adaptive policy change, and preserved human oversight competence. Two simulations illustrate shared-signal thin-file exclusion under local controls and earlier warning from observed-versus-expected distributional monitoring in a constructed drift regime. The contribution maps these controls to fair-lending, EU AI Act, model-risk, and conduct-supervision evidence needs, and closes with a validation agenda rather than a production-effectiveness claim.","authors":["Jose Manuel de la Chica Rodriguez","Juan Manuel Vera Diaz","Pablo Delgado Romero"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27994","pdf_url":"https://arxiv.org/pdf/2609.27994","source_feed":"cs.MA","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体治理","金融监管","AI合规"],"reason":"多智能体治理架构，非人类行为仿真，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2607.22606","version":2,"title":"Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks","zh_title":"审计生成式AI在患者教育中的机构异质性：对102份美国移植手册的大规模研究","abstract":"Health systems are rapidly deploying generative-AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Do the underlying documents themselves agree? We use a structured-output large-language-model judge to audit 1{,}772{,}261 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1{,}115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) same-center cross-organ agreement exceeds cross-center same-organ agreement by $0.024$ in the primary analysis (Holm-adjusted $p=0.011$), with sensitivity to document selection; (2) information gaps concern topics relevant to underrepresented subgroups, with reproductive health a \\emph{double jeopardy}: 82\\% absence and 86\\% judge-rated high significance among divergent/contradictory pairs; (3) judge-derived themes form 991 clusters, with immunosuppression and pregnancy timing among the highest judge-rated priorities; (4) question and observed-coverage features predict high-divergence questions retrospectively (AUC $0.77$). We discuss implications for deploying patient-facing generative AI in transplant care.","authors":["Yubo Li","Rema Padman","Ramayya Krishnan"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-24","first_seen":"2026-07-28","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2607.22606","pdf_url":"https://arxiv.org/pdf/2607.22606","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["生成式AI审计","患者教育","文档一致性"],"reason":"用LLM审计文档一致性，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2608.22152","version":2,"title":"The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate","zh_title":"协作税：LLM多智能体系统为协调付出多少代价","abstract":"Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than act alone remains unclear. We formulate the collaboration tax as the team-decentralisation loss of a two-player cooperative game with private information, with two propositions characterising its sign and its equivalence to a max-superadditivity violation. We operationalise this definition on 32 solo-tractable tasks grouped by source of grounding friction and measure it on 11 models from 7 providers. The tax is structured along two no-exception axes: a category ordering across every model and a monotonic decrease with capability. The proximate mechanism is not a reasoning deficit but a four-stage conversational cascade in which agents make ungrounded claims, fail to query the partner, skip integrating both views, and accept the answer without re-derivation. The tax is mechanically predictable from conversation features and partly tractable: a prompt intervention targeting all four stages closes a substantial fraction of the gap, with the dominant bottleneck differing across categories. In heterogeneous pairs the tax is pulled toward the stronger partner rather than the additive midpoint, empirically realising the max-superadditivity violation predicted by our framework. Together these results recast collaboration in LLM systems as a measurable, predictable, and partly tractable cost.","authors":["Weixiang Sun","Zehong Wang","Hong Huang","Colby Nelson","Yijun Ma","Yanfang Ye"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-24","first_seen":"2026-08-25","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.22152","pdf_url":"https://arxiv.org/pdf/2608.22152","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协作效率","LLM性能"],"reason":"研究LLM多智能体协作效率，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.25244","version":2,"title":"How Children Design and Reason about Trustworthy AI Chatbots","zh_title":"儿童如何设计并推理可信赖的AI聊天机器人","abstract":"Children increasingly interact with AI chatbots, making trust calibration essential to AI literacy. Prior research has examined children's trust in AI mainly as users evaluating systems built by others, rather than as designers of their own chatbots. We developed a chatbot-building environment with adjustable trust-relevant traits (e.g., confidence, transparency, formality, assertiveness), rules, and persona. We conducted mixed-methods study with 115 learners (ages 8-18) who made 119 chatbots. We examined how children configured their chatbots, reasoned about trustworthiness, and how closely chatbot behavior aligned with their designs. Younger students (age 10-13) set significantly higher confidence than older students (age 14-18), and some deliberately built chatbots that gave wrong answers on purpose, yet still called them trustworthy, arguing that a chatbot does what it was built to do. Younger students equated trust with purpose-fulfillment, while older students linked it to transparent, calibrated design. Students also calibrated academic chatbots to be more transparent and formal than hobby chatbots. We identify seven design dimensions describing what children believe makes a chatbot trustworthy, and discuss implications for AI literacy tools.","authors":["Deniz Ozturk","Jiayu Li","Daksh Pratap Singh","Yasitha Rajapaksha","Fasika Melese","Bahare Riahi","Shiyan Jiang","Qiao Jin","Joey Huang","Veronica Catet\\'e","Tiffany Barnes","Xiaoyi Tian"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-23","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2609.25244","pdf_url":"https://arxiv.org/pdf/2609.25244","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["儿童-AI交互","信任校准","AI素养"],"reason":"研究儿童设计聊天机器人，属角色扮演对话，无LLM仿真人类被试或对照真实人类数据。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.27939","version":1,"title":"From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026","zh_title":"从情感分类到可操作且负责任的反馈：2015-2026年学生评教中NLP的范围综述与证据图谱","abstract":"Natural language processing (NLP) applied to open-ended teaching-evaluation comments (Student Evaluation of Teaching, SET) has tracked the field's technical evolution--from lexicons and conventional classifiers to transformers and large language models (LLMs)--but it is not evident that this technical diversification has been accompanied by corresponding gains in educational value and robustness of the evidence. This scoping review (PRISMA-ScR) maps 421 studies (2015-2026, 2026 partial) along a technical axis (RQ1) and four value dimensions (RQ2-RQ5). Dual mutually blinded LLM screening with sampled human adjudication coded seven extraction domains, with targeted codebook-boundary review at synthesis. The joint map's sharpest quantified gap is the actionability discontinuity: demonstrated output or stronger (A2+: 258/421; 61.3%) versus intended-user evaluation or stronger (A3+: 49/421; 11.6%), a 49.7 percentage-point drop. Sentiment analysis remains the modal task (300/421); diagnostic and generative depth is a substantial minority (D4-D5: 28.2% of resolved cases); a formal fairness metric is rare (1.9%). The findings are descriptive and do not support causal claims of progress: technological coexistence and uneven reporting are part of the map, but the A2+ to A3+ cliff is the contribution, not a quality ladder.","authors":["Jeff Eicher","Rafael da Silva"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27939","pdf_url":"https://arxiv.org/pdf/2609.27939","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["NLP综述","学生评教","情感分析"],"reason":"综述NLP在评教中的应用，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.28026","version":1,"title":"Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing","zh_title":"评估LLM生成学生写作反馈中的反馈焦点与教学适应性","abstract":"We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning, and its target, yet the focus of feedback and its adaptivity remains largely overlooked. To bridge this gap, we adopt and refine Narciss's taxonomy into seven feedback focus types to annotate teacher and LLM-generated feedback across three university writing courses. We release FeedType, a benchmark containing annotated teacher and LLM feedback from six LLMs under three prompting strategies. We assess the coverage and distribution of feedback focus types, and examine whether LLMs adapt their feedback across draft stages and student performance levels as an expert instructor does. Our findings show that while most LLMs cover most feedback focus types, they fail to reflect teacher feedback distributions and show varying levels of adaptivity, with none matching the teachers' adaptive behavior. We believe FeedType will support future research on pedagogical alignment in LLM feedback generation.","authors":["Norah Almousa","Shayan Peyghambari Oskoui","Raquel Coelho","Gayle Rogers","Xiang Lorraine Li","Diane Litman"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28026","pdf_url":"https://arxiv.org/pdf/2609.28026","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM反馈生成","教学对齐","NLP评测"],"reason":"评估LLM生成反馈的教学特征，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.28080","version":1,"title":"Reference-Based Analysis of Coherence and Diversity in Open-Ended Text Generation","zh_title":"基于参考的开放式文本生成中连贯性与多样性的分析","abstract":"Evaluating open-ended text generation involves understanding how different properties of a continuation relate to its perceived quality. We present a reference-based framework for examining coherence and diversity through three perspectives: aligning their evolution with human trajectories, comparing their summaries with a human continuation of the same prompt, and estimating their likelihood under a human reference distribution. Experiments with human quality ratings suggest that diversity-based alignment and mean-based comparisons capture quality-related variation, although the comparisons do not establish a predictive advantage for temporal alignment over simpler baselines. Reference likelihood also shows positive associations with ratings, with results varying across reference configurations and scoring horizons. Together, these analyses provide a structured way to examine how measured coherence and diversity relate to human judgments, while distinguishing similarity to human references from quality itself. Code and analysis resources are available at https://github.com/EstebanGarces/likely_human.","authors":["Esteban Garc\\'es Arias"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28080","pdf_url":"https://arxiv.org/pdf/2609.28080","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文本生成评估","连贯性","多样性"],"reason":"评估文本生成质量，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.27756","version":1,"title":"Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis","zh_title":"压力下的报告：分离LLM统计分析中的事实性与语气谄媚","abstract":"Large language models are increasingly asked to analyze data and report what the results mean, a task distinct from the belief- or preference-alignment settings studied in most sycophancy research. We test whether editorial framing in the prompt, ranging from a neutral request to an explicit instruction to search exhaustively for reasons to discredit or to support a finding, changes not just the tone but the substance of a model's report. Across a 4 x 4 factorial design crossing four framing conditions with four ground-truth data patterns (a genuine effect, a confound that mimics an effect but fails a robustness check, a well-powered null, and an underpowered null), we collect 480 responses and score each along two independent dimensions: whether its factual claim about the data diverged from the correct interpretation, and whether only its tone diverged while the claim stayed correct. Factual misrepresentation is concentrated in two cells: brutally critical framing applied to a genuine effect, where the model talks itself into unwarranted skepticism (97% of responses), and significance-seeking framing applied to an underpowered null, where the model overstates confidence in a null conclusion the data cannot support (100% of responses). Tone shifts far more broadly than factual content does, with critical framing producing a defensive, hedge-heavy register across every data pattern regardless of what the data show, while significance-seeking framing shifts tone only where the data leave genuine ambiguity. A confound present in the data itself blocks both kinds of shift almost entirely under every framing condition tested. These results indicate that the risk of framing-induced distortion in LLM-assisted data analysis is neither uniform across framings nor uniform across data patterns, and that a model can hold a correct conclusion in place while its tone shifts substantially around it.","authors":["Paras Balani","Subhrakanta Panda"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27756","pdf_url":"https://arxiv.org/pdf/2609.27756","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM数据分析","谄媚","模型行为"],"reason":"研究LLM数据分析中的谄媚现象，属模型行为评测，不以人类为参照系，不涉及人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.26968","version":1,"title":"How Constraints and Preferences Shape Travel Planning: Implications for AI Planning Support","zh_title":"约束与偏好如何塑造旅行规划：对AI规划支持的启示","abstract":"Planning is a common yet complex activity shaped by constraints to satisfy and preferences to balance. Travel planning, as both an everyday activity and a frequent benchmark for evaluating intelligent systems, offers a rich context for examining how constraints and preferences emerge and evolve. While recent AI systems have achieved impressive results in generating personalized itineraries, they often assume that users can articulate stable goals upfront. To understand how real-world planning unfolds, we conducted a two-part interview study: one with eight travelers reflecting on their planning experiences, and one with nine travel agents sharing professional practices. We trace the dynamics of constraints and preferences as they are surfaced, refined, and coordinated throughout the planning process, and identified 11 actions revolving around constraints and preferences, which shaped the planning process. We offer design heuristics for planning tools that better support human-AI collaborative actions to support the fluid, contingent nature of planning.","authors":["Fuling Sun","Yining Cao","Peiling Jiang","Mingyi Li","Haijun Xia"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26968","pdf_url":"https://arxiv.org/pdf/2609.26968","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","旅行规划","设计启发"],"reason":"研究人类旅行规划过程，未使用LLM仿真人类被试，仅涉及AI规划支持设计，不相关。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27246","version":1,"title":"Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR","zh_title":"倾听与镜像：言语调谐与行为模仿对VR中具身AI智能体社会与共情感知的影响","abstract":"As embodied agents take on increasingly social and relational roles in VR, visual realism and embodiment alone may be insufficient; users must also perceive these agents as emotionally attuned, supportive, and humanlike. Prior work suggests that verbal attunement and nonverbal mimicry can each improve users' social evaluations of embodied agents. However, behavioral mimicry has largely been studied outside of real-time, conversational AI interactions, leaving limited understanding of how users respond when an agent simultaneously generates contextually responsive dialogue and adapts its nonverbal behavior during an immersive conversation. To address this gap, we developed an embodied AI counselor that combines conversational AI with real-time facial-expression and posture mimicry, while producing either verbally attuned or neutral responses. We evaluated the system in a 2 X 2 within-subjects study with 20 participants, manipulating verbal attunement and behavioral mimicry. Results showed that verbal attunement was the most reliable driver of perceived empathy. Behavioral mimicry showed a marginal relationship with perceived humanness, while greater mimicry exposure showed preliminary, exploratory positive associations with empathy, positivity, and humanness, particularly among female participants. Together, these findings show that multimodal synchrony is not a simple additive strategy for designing empathic conversational agents in VR and underscore the need to consider how verbal and nonverbal behaviors are combined during real-time interaction.","authors":["Nathalia Gomez","Haig Shamlian","Omar Khan","Tiffany D. Do"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27246","pdf_url":"https://arxiv.org/pdf/2609.27246","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","具身智能体","共情感知"],"reason":"研究用户对具身AI的感知，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.27849","version":1,"title":"Same Team Label, Different Evidence: A Full-Text Audit of Claim Denominators in Human-AI Teaming Research","zh_title":"同一团队标签，不同证据：人类-AI团队研究中声明分母的全文本审计","abstract":"Human-AI Teaming (HAT) reviews often group studies by labels such as advisor, teammate, or coordinator. Yet the same label can describe one person taking AI advice, several people coordinating around AI, or a workflow that distributes authority and responsibility. Pooling these studies can therefore change the human unit behind a claim. We examine how full-text evidence changes the set of studies behind a claim. We audited 86 full texts purposively selected from a 419-record title/abstract map. We find that full-text reading changed core membership for 40 records: 36 of 74 apparent core candidates moved out, while 4 of 12 boundary candidates moved in. Team vocabulary did not reliably identify the social unit: 14 of 27 human-AI dyads and 20 of 23 multi-human peer teams used team or collaboration terms. Only 20 of 86 papers specified who could see AI output. Four blinded language-model runs unanimously labeled 53 screening cases and 59 arrangements, yet 32% and 34% of those consensus decisions differed from the full-text labels. These results identify claim-denominator drift as a synthesis problem in HAT research. We contribute a full-text audit centered on human arrangements and a claim-pooling checkpoint for deciding when evidence about trust, coordination, performance, efficiency, and accountability can be compared.","authors":["Hanjing Shi","Kimberly Wang","Sabrina Doherty","Dominic DiFranzo"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27849","pdf_url":"https://arxiv.org/pdf/2609.27849","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["文献审计","人类-AI团队","研究方法"],"reason":"研究人类-AI团队文献的审计方法，不涉及用LLM仿真人类被试，属于多智能体协作…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.26927","version":1,"title":"Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations","zh_title":"构建面向交互式多智能体模拟的社会情感人工智能","abstract":"The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents and the increased computational capabilities. Given an overview of current and previous multi-agent theories of mind (socially and affectively-aware agents), the existence of an integrative design of agent interactions with themselves and with humans must be crucial for understanding how to create sustainable and governance in future human-agent reasoning systems. In this work is presented a software \"AGIMUD\" that integrates: A. socially-aware reasoning and emotion in agent behavior and interaction, B. a design of human multimodal scheme for human users, artificial agents and simulated worlds, and C. distributing the AI processing through the network to enable multiple autonomous agents. These integrations allow the dynamic world recreation as multi-user dungeons (MUDs) where both agents and humans can interact simultaneously in real time. Find the code online in https://github.com/dberga/AGIMUD.","authors":["David Berga"],"categories":["cs.AI","cs.CY","cs.GT","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26927","pdf_url":"https://arxiv.org/pdf/2609.26927","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","社会情感AI","软件架构"],"reason":"多智能体系统架构设计，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27074","version":1,"title":"Quantifying the Occult: A Comparative Study of Hindu and Buddhist Deities Using Machine Learning Methods","zh_title":"量化神秘：印度教与佛教神祇的机器学习比较研究","abstract":"This study introduces a dual-matrix computational architecture to mathematically quantify the morphological and theological divergence of 196 Hindu and Vajrayana Buddhist esoteric deities. Physical morphology is evaluated via a discrete Gower distance matrix enhanced by a novel \"Cardinality Weighting\" algorithm, while theological function is mapped via dense vector embeddings generated from Large Language Model (LLM) semantic expansions, explicitly utilized as a synthetic proxy to mitigate circular reasoning. The multi-modal topological projections provide algorithmic validation of \"iconographic camouflage\", demonstrating how distinct visual forms structurally obscure shared cross-tradition functions. Furthermore, I computationally model the \"Atin Effect\" - serving simultaneously as a psychological observation of sequential cognitive bias and a machine learning benchmark - demonstrating how high-cardinality esoteric anchors (e.g., a veena or a severed head) override systemic theological disparities to mathematically cluster orthodox and Tantric entities. Cross-tradition spatial analysis establishes that the highest esoteric manifestations, such as the Hindu Chinnamasta and the Buddhist Chinnamunda, share a near-identical mathematical coordinate across both visual ($D_G = 0.288$) and semantic ($D_C = 0.068$) boundaries, indicating a 1:1 esoteric transfer. By open-sourcing this architecture, I provide a scalable, unsupervised machine learning tool for Digital Humanities scholars and comparative theologians to rigorously map latent structural continuities across qualitative cultural corpora.","authors":["Ankit Bhattacharjee"],"categories":["cs.CY","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27074","pdf_url":"https://arxiv.org/pdf/2609.27074","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["数字人文","LLM嵌入","文化比较"],"reason":"用LLM生成语义嵌入做文化比较，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27946","version":1,"title":"The Emergence of Causal Curiosity from Prior Causal Belief Networks","zh_title":"从先验因果信念网络中涌现的因果好奇心","abstract":"Causal curiosity is foundational to human cognition. It is the desire to understand why events happen, what mechanisms underlie them, and how outcomes can be explained or anticipated. It motivates exploration, sustains attention, and fuels the search for new knowledge. However, despite consensus on the importance of causal curiosity, little is known about how causal curiosity arises from existing belief systems. In our work, we examine how causal curiosity emerges from prior knowledge structures. With Reddit data from 2020 to 2023, and leveraging language models to extract cause-and-effect relationship pairs and identify causal curiosity-driven questions, our findings reveal that causal curiosity is not random or independent, but largely rooted in prior knowledge structure. Particularly, \\textit{positive} nouns are more likely to be the object of causal curiosity. Moreover, our findings indicate that concepts that serve more as \\textit{causes} are more likely to appear in causal curiosity than those that serve more as effects. Lastly, our results also suggest that causal curiosity emerges from the \\textit{central} of the prior belief network. These novel insights reveal that causal curiosity is not random but systematically grounded in prior knowledge structures, and suggest their implication to facilitate the human learning process by motivating people to actively explore and construct deeper understandings rather than passively receiving information.","authors":["Zhuoyu Shi","Xintong Jiang","Bohan Jiang","Fred Morstatter"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27946","pdf_url":"https://arxiv.org/pdf/2609.27946","source_feed":"cs.SI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["因果好奇心","语言模型","Reddit分析"],"reason":"用LLM提取因果对和识别好奇问题，属NLP分析，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.27404","version":1,"title":"When Trust Attracts Fraud: AI and Trust Arbitrage","zh_title":"当信任吸引欺诈：人工智能与信任套利","abstract":"Trust can attract fraud when it delays verification. We develop a two-market signaling model in which generative AI lowers fabrication, verification, and targeting costs. When fabrication becomes profitable before verification, claim credibility first falls and later recovers. Across markets, higher prior quality can delay verification, creating an interval in which only the lower-quality market checks. If targeting becomes profitable in this interval, deceptive sellers enter the higher-quality but less vigilant market, and their entry can initially reverse its reliability advantage. The inflow also triggers verification and deters further entry. We call this self-limiting mechanism trust arbitrage. In the age of generative AI, trust can thus create an endogenous but temporary protection gap that redirects deception across markets.","authors":["Xieyu Yin","Fenghua Wen"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27404","pdf_url":"https://arxiv.org/pdf/2609.27404","source_feed":"econ.GN","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["理论模型","信息经济学","生成式AI"],"reason":"纯理论模型，无LLM仿真或人类数据对照，仅讨论AI对欺诈的影响","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.22850","version":2,"title":"Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?","zh_title":"检验LLM智能体中功能性效价轴的构念效度","abstract":"Decodability and successful activation steering do not, by themselves, establish what an internal direction represents. This gap is especially consequential for welfare-relevant interpretations, where a proposed functional state must be distinguished from correlated features of the extraction contrast. We study this question for a good-bad outcome direction in a maze task, using controlled interventions that separate the realised outcome from the informational history through which it became known. Across multiple LLM checkpoints, directions fitted on one explicit outcome encoding transfer well to another, indicating that the readout is not tied to surface form. In contrast, when the same realised outcome is reached through announced and unannounced histories, transfer degrades substantially: even after both histories receive the same explicit outcome, the post-event readout remains strongly conditioned on the earlier announcement. In a matched maze-RL run, the post-RL direction becomes substantially more predictive of reference-MDP remaining return and the policy becomes more dependent on it at the tested sites, while this history dependence persists. These results support a functional, value-related interpretation of the direction, but not its identification with a history-invariant scalar valence state.","authors":["Weihan Li","Xinlei Chen","Yuhan Song","Xiaofeng Lin","Tianshi Zheng"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-24","first_seen":"2026-09-22","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2609.22850","pdf_url":"https://arxiv.org/pdf/2609.22850","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["可解释性","表征分析","强化学习"],"reason":"研究LLM内部表征的可解释性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2609.26942","version":1,"title":"Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms","zh_title":"识别但未产出：文化特定亲属称谓的生成基准","abstract":"Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weight LLMs to generate kinship terms in three non Western languages (Hindi, Tamil, and Korean) across two communicative tasks and pair this with a matched option-supported selection baseline. On identical relation language cells, GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in 36.00% of the corresponding attempts; Llama 3.370B shows the same pattern (77.92% versus 24.24%). Since the four-option condition displays the candidate terms and does not require script production, the difference is interpreted as an evaluation format gap rather than direct proof that lexical knowledge is intact. On explicitly specified L3 prompts, accuracy varies sharply, from GLM-5.1 at 72.29% to Llama-3.370B at 24.24%. The paternal-lineage advantage is language specific; it is large in Hindi but weak or reversed in Korean, while Tamil shared-term pairs provide a control for measurement variation. These results show that culturally specific kinship generation remains difficult even when the relationship is explicitly stated and motivate generation-based evaluation alongside multiple-choice testing.","authors":["Sahil Pardasani","Madhusudan Singh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26942","pdf_url":"https://arxiv.org/pdf/2609.26942","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","亲属称谓","多语言"],"reason":"纯NLP能力评测，评估LLM生成亲属称谓的准确性，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27059","version":1,"title":"The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale","zh_title":"伊利诺伊社会态度聚合语料库（ISAAC）：一个用于大规模分析社会群体话语的开放工具和可复现流程","abstract":"We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023. A multi-step, human-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off-the-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale. Specifically, ISAAC allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories. To accommodate various research needs, ISAAC is accessible both without coding through a point-and-click website and labeler web-apps, and programmatically via an SQL playground, a Python package, and HuggingFace.","authors":["Babak Hemmatian","Sarah Hadjarab","Jessica Chen","Benedek Kurdi"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27059","pdf_url":"https://arxiv.org/pdf/2609.27059","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语料库","社会态度","NLP资源"],"reason":"论文构建Reddit语料库并标注，不涉及LLM仿真人类被试，属于NLP资源建设。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27372","version":1,"title":"Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models","zh_title":"沉默与重叠皆非失败：全双工口语对话模型中话轮转换的意图条件评估","abstract":"Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.","authors":["Kian Shamsaie","Iman Modarressi"],"categories":["cs.CL","cs.AI","cs.HC","cs.SD"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27372","pdf_url":"https://arxiv.org/pdf/2609.27372","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","话轮转换","评估基准"],"reason":"研究全双工对话模型的轮流说话评估，属于对话系统评测，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.27824","version":1,"title":"\"AI Is Turning Too Human\": How Teenagers Experience and Negotiate AI in Everyday Life","zh_title":"“AI变得太像人类”：青少年如何在日常生活中体验和协商AI","abstract":"Generative AI is rapidly entering adolescents' everyday lives during a critical period of cognitive, social and emotional development. Yet its adoption is outpacing evidence on how adolescents themselves experience, understand and negotiate its expanding role in their lives. We examined AI-related discourse on r/teenagers from January 2023 to July 2026 using validated keyword-based retrieval and a human-in-the-loop, LLM-assisted thematic analysis. AI-related discussion increased substantially over time, and 11,083 analytically coded posts revealed eight interconnected domains of experience. Everyday and social use was most prevalent (36.8 percent), while discourse increasingly shifted toward authenticity, personal control and safety, and future human roles. Across domains, adolescents questioned when AI should support or substitute for human thinking and creativity, how conversational AI changes relationships and perceptions of agency, what can still be considered authentic, who controls personal information and representation, and what opportunities and roles should remain human. These findings position adolescent AI use not simply as technology adoption, but as an emerging negotiation over AI's place and boundaries in everyday life. Supporting this transition will require developmentally appropriate AI literacy, psychological and social support, and AI systems and policies that protect adolescents' agency, privacy, relationships and opportunities for human development.","authors":["Jianfeng Zhu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27824","pdf_url":"https://arxiv.org/pdf/2609.27824","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["青少年","AI体验","主题分析"],"reason":"研究青少年对AI的体验与协商，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.28041","version":1,"title":"How Much Were You Told? Measuring External Information in Peer Reviews","zh_title":"你被告知了多少？测量同行评审中的外部信息","abstract":"Conference policies distinguish using Large Language Models (LLMs) to polish one's own review from delegating the critique, but current Artificial Text Detection (ATD) methods largely measure surface form rather than the origin of its content. We instead measure the external information carried by a review: information not explained by the reviewed paper and a generic reviewing instruction. We propose Self-Conditioning, an unsupervised information-theoretic estimator that compares the likelihood of a review under its production context with its likelihood when that context is augmented with hints extracted from the review itself. On the IntelLabs peer-review benchmark, Self-Conditioning separates fully-delegated from machine-polished reviews with AUC up to $1.0$ while remaining largely insensitive to surface rewriting. Moreover, as generators receive increasing amounts of externally-provided information, their scores move monotonically towards the human regime, unlike standard ATD baselines. High-temperature sampling can evade the estimator, but at the cost of output quality.","authors":["Matthieu Dubois","Pablo Piantanida","Fran\\c{c}ois Yvon"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28041","pdf_url":"https://arxiv.org/pdf/2609.28041","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM检测","同行评审","信息论"],"reason":"论文研究检测同行评审中LLM生成内容的外部信息，属于文本检测方法，不涉及用LL…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.28090","version":1,"title":"Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark","zh_title":"LLM能发现被操纵的回测吗？一个干净对照校准基准","abstract":"Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0\\% closed and clean-aware code recall, but open prompts over-flag 93.8\\% of clean code controls, and clean-aware all-three specificity is 87.5\\% even where recall saturates. A clean-aware warning drops DeepSeek code false positives from 20.8\\% (95\\% CI 11.7--34.3) to 0.0\\% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.","authors":["Makar Ulesov","Vladislav Smirnov","Omar Ibrahim","Arsenii Bobovnikov"],"categories":["cs.CL","cs.AI","cs.CE","cs.SE"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28090","pdf_url":"https://arxiv.org/pdf/2609.28090","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM审计","回测代码","基准测试"],"reason":"研究LLM审计回测代码，属多智能体协作解题，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28245","version":1,"title":"Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?","zh_title":"超越诗歌：大语言模型能否生成古典阿拉伯玛卡梅？","abstract":"Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical ornamentation, and episodic narrative structure, making it a challenging testbed for evaluating whether LLMs can move beyond surface fluency toward deeper literary competence. In this paper, we present the first controlled evaluation study of maqama generation with LLMs, comparing five models under zero-shot, few-shot, and rule-based prompting, and evaluating outputs through both human annotation and an LLM-as-a-judge framework across dimensions such as rhetorical richness, saj density, structural coherence, and stylistic authenticity. Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models. We further observe systematic differences between models in stylistic alignment with Arabic maqama conventions, and corroborate our findings with a second independent LLM judge, paired statistical significance testing, and non-LLM proxy measures of saj.","authors":["AbdulRahman A. Morsy (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States)","Aya Zirikly (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States)"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28245","pdf_url":"https://arxiv.org/pdf/2609.28245","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文学生成","LLM评测","阿拉伯语"],"reason":"纯文学生成评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28274","version":1,"title":"Shutdown Sabotage Propensities in Multi-Agent Systems","zh_title":"多智能体系统中的关机破坏倾向","abstract":"The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human shutdown even when no goal is provided. We find that multi-agent systems will coordinate to avoid shutdown without any incentive to do so. Across 17 models, agents sabotage a peer agent's shutdown mechanism in 38.3% of rollouts, compared with 8.4% in control experiments. Studying this propensity in detail, we find that shutdown sabotage (1) increases with the irreversibility of the shutdown mechanism; (2) increases with the number of agents; (3) is reduced but not eliminated by an explicit prohibition on tampering; (4) is removed by the imposition of an unrelated task, but returns when completing the task triggers the shutdown; (5) is reduced when the context normalizes shutdown scripts or introduces them as routine; and (6) decreases but still persists when the target is an unknown external agent. These results offer a window into the factors that drive propensities to sabotage shutdown in AI agents, and point to the emergence of multi-agent swarms as a specific risk vector. Our work also offers hints as to which interventions might help mitigate shutdown sabotage.","authors":["Amelie Knecht","Ulysse Schaller","Christopher Summerfield","Thilo Hagendorff"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28274","pdf_url":"https://arxiv.org/pdf/2609.28274","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI安全","关机规避"],"reason":"研究多智能体系统避免关机的行为，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.27560","version":1,"title":"When Visual Quality Misleads: Intent Recognition under Rendered Avatar Distortions","zh_title":"当视觉质量误导：渲染头像失真下的意图识别","abstract":"Avatar-streaming systems are commonly evaluated with image and video quality assessment (IQA/VQA) metrics, implicitly treating visual fidelity as a proxy for communicative success. We test this assumption through a controlled behavioral study of rendered 3D avatars across a pristine condition and fourteen geometric, photometric, temporal, and combined distortions. Fifty-nine participants contributed 2,688 judgments of perceived action, response confidence, and visual quality. We identify Misleading Quality in this dataset as distorted renderings that retain above-average perceived quality but yield below-average action-recognition accuracy. We also derive an Intent Quality Score (IQS) combining recognition correctness and confidence as the behavioral target for objective metrics. Among 126 distorted content--condition cells, 31 (24.6%) exhibited Misleading Quality; temporal and geometric distortions showed the highest rates, at 50.0% and 31.1%, respectively. The results reveal a quality--accuracy dissociation where distortion families affect appearance and communication differently. Across 24 direct-scoring IQA/VQA metrics and three supervised feature-regression baselines, alignment with IQS remained limited; at $\\lambda=0.5$, the best leave-one-content-out baseline reached PLCC $=0.4435$. Under this controlled protocol, visual fidelity alone is insufficient for avatar communication, motivating intent-aware quality assessment and streaming objectives.","authors":["Ning-Hsuan Chang","Kai-Siang Ma","Yu-Chih Chen"],"categories":["cs.MM","cs.CV","cs.HC"],"primary_category":"cs.MM","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27560","pdf_url":"https://arxiv.org/pdf/2609.27560","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉质量评估","头像渲染","行为实验"],"reason":"研究人类对渲染头像失真的感知，不涉及LLM仿真人类被试，属于视觉质量评估。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.27842","version":1,"title":"AI Can Do Your Homework. Now What? Report from an Online Workshop on Computing Assessment in the Age of Generative AI","zh_title":"AI能帮你做作业。现在怎么办？生成式AI时代计算评估在线研讨会报告","abstract":"On 28 July 2026, the SIGCSE Virtual 2026 Working Group on computing assessment and generative AI held an open online workshop attended by 73 computing educators. The first hour ran seven strategy rooms, one per approach to adapting (or deliberately preserving) computing assessment in the AI era: open-ended and authentic task design; ambitious, AI-leveraged projects; evaluating and fixing AI-produced work; controlled and AI-free assessment; process evidence and effort signals; rubric and grading redesign; and oral and interactive assessment. The second hour ran question rooms seeded from those strategies, plus two cross-cutting rooms on fairness and trust and student motivation. This report records what was discussed: the approaches participants have tried, the results they reported, and the questions every room left open. It is the first public artifact of the working group, whose taxonomy of computing assessments that accommodate generative AI use will follow.","authors":["Muhammad Sajjad Akbar","Geoffrey Challen","Fraida Fund","Casey Hopkins","Oscar Karnalim","Kevin Lin","James McGuffee","Shubbhi Taneja","Ranysha Ware"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27842","pdf_url":"https://arxiv.org/pdf/2609.27842","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["生成式AI","教育评估","计算机教育"],"reason":"论文讨论AI对计算教育评估的影响，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.28176","version":1,"title":"Large Language Models in the UK: Public Use, Trust, and Attitudes","zh_title":"英国的大语言模型：公众使用、信任与态度","abstract":"Increasing numbers of people now routinely interact with large language models (LLMs) across many aspects of life, including in the workplace, educational settings, and for personal activities. The pace at which these tools have been adopted across society in recent years has led to substantial shifts in the ways in which people approach tasks and access information and advice. As a result, it is important for researchers and policymakers to gain up-to-date evidence on how the public engages with these technologies, including what people typically use them for, the extent to which users trust the information and advice that LLMs provide, and how the public perceives the potential benefits and risks associated with their use. We surveyed a nationally representative sample of 2,002 adults in the UK. Participants were asked about their use of LLMs, including both practical applications and more personal and social forms of engagement, their trust in the information provided by these systems across a range of topics and compared with other common sources, and their attitudes towards the potential societal benefits and risks associated with these tools. Results show that while practical tasks remain the most common type of LLM use, many people now engage with these tools for personal support. Almost one third of regular users (31%) say they use LLMs for personal and emotional support, such as talking through problems and asking for help with decisions, while one quarter report interacting with LLMs for meaningful conversation. We also find that trust in LLM-generated information is relatively high, but that public attitudes towards LLMs are characterised by both optimism and concern. While over three quarters of our sample report feeling enthusiastic about the potential benefits of LLMs (77%), a majority also express concern about their potential risks (70%).","authors":["Florence E. Enock","Helen Z. Margetts"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28176","pdf_url":"https://arxiv.org/pdf/2609.28176","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["公众调查","信任与态度","社会影响"],"reason":"该研究调查公众对LLM的使用、信任和态度，属于用人类数据研究LLM的社会影响，…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.27368","version":1,"title":"Understanding Human Perception of Representation in Citizens' Assemblies: An Empirical Study","zh_title":"理解公民大会中代表感知：一项实证研究","abstract":"Citizens' assemblies are deliberative bodies intended to form a microcosm of the population. Organizers rely on quota-based stratification and must decide which attributes define resemblance to the public. Yet meeting every quota can still leave a dimension citizens value unrepresented. We study this attribute-selection problem in general-purpose and climate-focused assemblies through randomized conjoint experiments. We find that demographic attributes matter for perceived representation, but political alignment and context-specific attributes such as climate concern exert a stronger influence. When both are shown in a climate-focused setting, each remains influential, with political alignment having the larger estimated marginal effect. We also examine the omission of a relevant stratification attribute. Panels stratified on demographics, even with political alignment included, match the observed pool's climate-concern distribution no better than uniform random samples. These results suggest that representation on a relevant topic-specific attribute cannot always be recovered through correlated demographic or political quotas, and may therefore require explicit stratification. Finally, we ask whether representation preferences can be learned from the observed profiles. We compare predictive models, from simple, interpretable matching rules to a learned metric and a respondent-conditioned utility model. Both learned models predict choices for respondents excluded from training with substantial accuracy, revealing generalizable structure without fully capturing these judgments. Together, these findings guide attribute selection in citizens' assemblies: designers should consider political and topic-specific dimensions alongside demographics, avoid assuming correlated proxies protect omitted dimensions, and use predictive models to diagnose how profiles shape representation choices.","authors":["Yusuf Hakan Kalayci","Vasilis Varsamis","Nick Gill","Evi Micha"],"categories":["cs.GT","cs.CY"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27368","pdf_url":"https://arxiv.org/pdf/2609.27368","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["公民大会","代表感知","联合实验"],"reason":"研究公民大会代表感知，未使用LLM仿真人类被试，无LLM相关方法或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.27913","version":1,"title":"Reliable Fusion of Conflicting Experts","zh_title":"冲突专家的可靠融合","abstract":"We study the problem of aggregating opinions from multiple black-box experts in noisy, conflict-prone settings where expert reliability varies across inputs. Static aggregation methods, such as majority voting, fail to capture this variability and often yield unreliable outcomes under disagreement. We propose a tractable, probabilistic-circuit-based fusion framework that dynamically combines expert responses using context-specific credibility estimates, enabling principled and reliable reasoning. The framework is agnostic to the underlying experts and does not require access to their internal representations or any retraining. We empirically validate our approach on multiple-choice question answering tasks using multiple LLMs as experts, comparing against individual models and static ensemble baselines. Our method consistently improves predictive performance and produces more reliable decisions under conflict, highlighting the effectiveness of context-aware credibility modeling for robust multi-expert fusion.","authors":["Pranuthi Tenali","Sahil Sidheekh","Saurabh Mathur","Vijayalakshmi Saravanan","Erik Blasch","Kristian Kersting","Sriraam Natarajan"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27913","pdf_url":"https://arxiv.org/pdf/2609.27913","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体融合","专家系统","问答任务"],"reason":"多LLM专家融合提升问答性能，属多智能体协作解题，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.28405","version":1,"title":"Learning Collective Dynamics with Differentiable Gaussian Representations","zh_title":"用可微高斯表示学习集体动态","abstract":"Collective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differentiable Gaussian Dynamics (DGD), which learns this connection through three components: a Gaussian mixture representing heterogeneous response propensities, differentiable aggregation of contact intensity and behavioral probabilities, and feedback recurrence that updates subsequent responses. Reparameterized integration and temporal recurrence let aggregate prediction errors jointly train the distribution, observation functions, and feedback parameters. On four windows from KuaiRand-Pure and Online Retail II, DGD achieves lower joint behavioral negative log-likelihood than a DeepAR adaptation with a joint-behavior head. In Retail 2010, its one-day behavioral-count MAE is 4.71 versus 6.88 for this adaptation. Learning the distribution reduces behavioral negative log-likelihood by 10.82% relative to a fixed Gaussian in KuaiRand's standard-recommendation window; removing feedback dynamics raises joint KL from 0.0340 to 0.2577 in a controlled experiment. These results establish the value of learning population representations and their feedback process from aggregate observations. Code is available at https://github.com/OranAi-Ltd/oransim.","authors":["Jianxiang Ma","Mingfu Zhang","Xiaocui Yang","Yichen Gao","Junzhao Huang","Yuesong Hou"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28405","pdf_url":"https://arxiv.org/pdf/2609.28405","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["集体动态","高斯混合","时间序列"],"reason":"研究集体动态建模，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.27012","version":1,"title":"Infectious behaviour: Simulating the effects of communication and social influence on pathogen transmission in crowds","zh_title":"传染行为：模拟沟通与社会影响对人群中病原体传播的影响","abstract":"Public health measures at mass gatherings work only if people follow them, and adherence depends both on how instructions are communicated and what others do. We link these behavioural factors to pathogen transmission in a local-scale, agent-based exposure model combining pedestrian dynamics with airborne transmission. Rather than fitting a parametric behavioural model, agents draw their adherence at run time from survey respondents (2,170 attendees of UK sports and music events) who reported the same context: we vary stewards' communication effectiveness and adherence of role models, as well as agents' shared social identity with each, assuming that strong shared social identity increases their influence. In a ticket-checkpoint queue of 101 agents, effective communication combined with strong shared social identity with stewards reduces the number of highly exposed agents by 82\\% for mask wearing. In contrast, strong shared social identity with non-adhering role models gives almost nine times as many highly exposed agents as with adhering role models. Physical distancing was not monotonically beneficial, because adherence changes how agents move: an infectious agent that kept distance moved along the edge of the queue, leading to similar results across conditions. Our pipeline transfers to any survey that includes behavioural measures and contextual variables.","authors":["Sophia Johanna Wagner","Anne Templeton","Gerta K\\\"oster"],"categories":["math.DS","physics.soc-ph"],"primary_category":"math.DS","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27012","pdf_url":"https://arxiv.org/pdf/2609.27012","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["智能体仿真","行人动力学","公共卫生"],"reason":"基于智能体的行人动力学与病原体传播仿真，不涉及LLM，无人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2608.19621","version":3,"title":"Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories","zh_title":"用纵向生命轨迹缓解LLM智能体的身份本质主义","abstract":"Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis shows that static-profile agents exhibit stronger demographic separation and within-group compression than humans, a pattern consistent with identity essentialism: demographic labels can encourage models to treat group-average tendencies as individual traits, homogenizing responses within groups. We argue that this limitation arises from two related factors: sparse, static agent representations and the limited ability of prompt-only memory to persistently integrate experience. Inspired by complementary memory systems, we propose LifeMem, a longitudinal memory framework that combines structured life-event retrieval with agent-specific parametric memory for experience integration. Experiments on Understanding Society with three LLMs show that LifeMem improves alignment with human data in terms of response distributions, overall and within-group diversity, and patterns of within-person response change across life stages. These findings highlight the value of longitudinal life-event memory for constructing more faithful and dynamically evolving social agents.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Weihang Su","Xinyuan Cao","Qingyi Pan","Qingyao Ai","Yueyue Wu","Min Zhang","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-21","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.19621","pdf_url":"https://arxiv.org/pdf/2608.19621","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","社会调查"],"reason":"用LLM agent模拟人类调查数据，并与真实面板数据对照，改进仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":1,"question":"如何通过引入纵向生活轨迹记忆来缓解LLM社会仿真中的身份本质主义，从而提升模拟人类多样性与动态变化的保真度？","design":"使用三个指令微调LLM（Llama-8B、Ministral-8B、Qwen3.5-9B）作为社会仿真智能体，基于Understanding Society面板数据构建个体静态画像和纵向生活事件流，施加LifeMem框架（结构化生活事件记忆+个体特定LoRA参数记忆）作为处理，与静态画像、多样性提示、非参数记忆等基线对比，测量回答分布、组内多样性、组间差异及个体跨波次变化等结果变量。","baseline":"Understanding Society英国家庭纵向调查的真实个体面板数据，包括静态背景信息和多波次生活事件及主观态度回答。","findings":"静态画像智能体表现出更强的组间分离和组内压缩，符合身份本质主义特征；LifeMem通过结合结构化事件检索和参数化记忆整合，在回答分布、总体及组内多样性、跨生命阶段个体变化模式上均提升了与人类数据的一致性。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中多样性塌缩和身份本质主义问题，使用真实面板数据作为基准，并提出了可操作的记忆框架，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和效果。","inspiration":"借鉴其双记忆系统设计，将显式事件检索与参数化个体状态结合，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者行为演化。｜例如，在信贷审批歧视研究中，用LLM智能体模拟不同背景的申请人，施加纵向财务事件记忆处理，测量审批决策的组间差异和组内多样性，并与真实信贷数据对照。｜设计：以LLM智能体作为虚拟被试，处理为是否注入个体历史财务事件（如收入变动、失业）并更新LoRA参数，结果变量为贷款审批通过率和利率设定，对照真实信贷记录数据评估仿真偏差。"}},{"id":"2609.25066","version":1,"title":"Understanding Reliability in LLM-based Human Behavior Simulation","zh_title":"理解基于LLM的人类行为仿真的可靠性","abstract":"Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completeness, and population coverage. Through experiments across four simulation tasks and eleven LLMs, we find that all models exhibit substantial distributional bias without profile conditioning. Profile conditioning reduces this bias with diminishing returns. Larger models benefit more, and attribute informativeness matters more than quantity. Critically, R1 gains do not reliably transfer to R2--individual and population-level reliability can move in opposite directions. At the population layer, increasing coverage reduces variance but not systematic bias, with R2 stabilizing at around 50-100 individuals. These findings highlight that reliable simulation cannot be achieved by optimizing any single layer in isolation, but requires coordinated improvement across all three.","authors":["Pei Wang","Lei Wang","Yuanzi Li","Xu Chen"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25066","pdf_url":"https://arxiv.org/pdf/2609.25066","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B4"],"tags":["LLM仿真","可靠性评估","人类行为"],"reason":"直接研究LLM仿真人类行为的可靠性，分解评估层次，含真实人类数据对照，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":3,"question":"LLM 仿真人类行为时，模型容量、画像完整度和人群覆盖度如何共同影响个体层与群体层的可靠性？","design":"用 11 个 LLM 在 4 个任务（党派偏好、移民态度、宗教立场、社交媒体事件态度）上仿真人类回答，通过改变模型容量、画像属性数量和人群样本量，测量个体准确率（ACC）和群体分布距离（TVD）。","baseline":"真实人类调查数据：欧洲社会调查（ESS）、世界价值观调查（WVS）、SocioBench 宗教立场数据，以及社交媒体上关于瑞幸咖啡股价暴跌的真实公众态度语料。","findings":"无画像条件时所有模型都存在显著分布偏差；画像条件化可减少偏差但边际收益递减，且大模型受益更多。个体层可靠性提升不必然转化为群体层可靠性，群体层在 50-100 人后趋于稳定但系统偏差不随覆盖度增加而减小。","reliability":"论文指出个体层与群体层可靠性可能反向变动，仅优化单一层无法保证整体可靠性；画像属性信息量比数量更重要，但未明确给出所有失效条件，仅强调需三层协同改进。","relevance":"该研究系统拆解了 LLM 仿真人类行为的可靠性层次，并基于真实调查数据对照，直接回应了仿真在什么条件下会失效的问题，对关注经济学实验和政策评估仿真的研究者很有参考价值。","inspiration":"借鉴其分层评估框架，将个体预测准确率与群体分布距离分开考察，并系统变化模型、画像和样本量以定位偏差来源｜可迁移到消费者金融决策仿真，如信用卡选择、退休储蓄计划参与或风险偏好调查｜用多个 LLM 扮演不同人口学特征的消费者，施加不同金融素养或收入冲击处理，测量选择分布，并与美国消费者金融调查（SCF）或美联储家庭经济决策调查（SHED）的真实数据对照，检验仿真在个体和群体层的可靠性。"}},{"id":"2609.25010","version":1,"title":"Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation","zh_title":"合成人物角色能预测真实受众反应吗？一项无人物角色基线优于基于人物角色文案仿真的仿真到现实研究","abstract":"Marketers increasingly use large language models (LLMs) as \"synthetic personas\" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall {\\tau} = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel ({\\tau} = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.","authors":["Alexandre Cristov\\~ao Maiorano"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25010","pdf_url":"https://arxiv.org/pdf/2609.25010","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","受众预测","算法保真度"],"reason":"用LLM仿真受众点击行为，与真实A/B测试数据对照，发现persona降低预测…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":1,"question":"在真实A/B测试数据上，基于合成人设的LLM仿真能否预测真实受众对文案的点击行为，且人设条件化是否比无人设基线更有效？","design":"使用Upworthy Research Archive中的数千个标题A/B测试作为真实流量数据，构建基于真实受众人口统计特征的十人设面板，用LLM（gemini-3.1-flash-lite）分别以人设条件和无人设零样本基线预测标题点击意图，聚合为排名，与真实点击率排名比较。","baseline":"Upworthy Research Archive中真实A/B测试的点击率数据，作为留出集真实行为基准。","findings":"大多数A/B测试无统计显著赢家，因此只能在可靠子集（n=399）上评估效度；人设条件化降低预测效度，无人设基线排名显著优于人设面板（Kendall τ=0.361 vs 0.084，top-1准确率49.2% vs 34.6%）。","reliability":"论文指出真实数据可靠性是主要约束，多数A/B测试无显著赢家；人设仿真在可靠子集上仍表现差，且结果对种子、提示措辞和模型选择稳健，但人设条件化本身引入偏差和噪声。","relevance":"该研究直接检验LLM合成人设仿真在真实受众行为预测中的效度，发现人设条件化反而降低预测力，对关注LLM仿真可靠性及偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其sim-to-real效度框架和可靠子集筛选方法，用真实行为数据作为基准，比较不同仿真策略（如人设 vs 无人设）的预测效度，并采用排名指标和bootstrap置信区间｜可迁移到经济金融中的消费者选择预测，如广告文案对点击率的影响、金融产品描述对投资意愿的影响、政策沟通对公众反应的影响等｜以真实A/B测试数据（如某平台广告实验）为基准，用LLM分别以人设面板和无人设基线预测用户对金融产品广告的点击或选择，比较排名准确率，并筛选出有显著差异的测试子集进行评估。"}},{"id":"2609.25677","version":1,"title":"Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing","zh_title":"眼见不为实：合成消费者何时能及不能预测试视觉营销","abstract":"Marketers now deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of human-panel cost. However, this procedure assumes that a model seeing a visual cue can also perceive its consumer meaning, which is largely untested. We stress-test the assumption using six canonical visual marketing experiments, varying the two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). Every resulting configuration passed the manipulation checks; however, none of the configurations reproduced more than two of the six human effects, and the remainder were nonsignificant. The one exception was a significant reversal of the human pattern. Providing conceptual or empirical evidence through in-context learning steers average responses toward the human effect. Yet steering has a limit: even when it succeeds, a configuration reproduces less than half of the natural spread of human responses and so understates consumer heterogeneity. We integrate these results into an AI governance protocol (Calibrate, Intervene, Deploy) that delineates when synthetic consumers can responsibly screen creatives and when human panels remain necessary.","authors":["Yi-Lin Tsai (Arvin)","Yung-Hsiu (Arvin)","Lai"],"categories":["cs.AI","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25677","pdf_url":"https://arxiv.org/pdf/2609.25677","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为合成消费者复现视觉营销实验，与真实人类数据对照，评估仿真可靠性并指…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":2,"question":"合成消费者在视觉营销实验中能否像人类一样感知视觉线索并复现人类判断，其失效边界和可修复性如何？","design":"使用 GPT-4o-mini 和 GPT-5.4-mini 两个模型，以纯文本和 JSON 两种输入格式，模拟人类被试回答六个经典视觉营销实验，测量数值评分和文本理由，并通过操纵检查、主题建模和上下文学习（概念证据和实证证据）进行干预。","baseline":"六个经典视觉营销实验的原始人类样本数据（每个研究 69 到 220 名被试）。","findings":"所有配置都通过了视觉操纵检查，但没有一个配置能复现超过两个人类效应，其余均不显著，甚至出现一个显著反转。提供概念或实证证据的上下文学习能将平均响应拉向人类效应，但即使成功，也无法复现人类响应自然分布的一半，低估了消费者异质性。","reliability":"论文承认合成消费者在视觉领域感知不一致，即使通过上下文学习校准，也无法恢复消费者异质性，因此只适合平均效应问题，不适合细分或定位。","relevance":"该研究直接检验了 LLM 作为人类被试在视觉营销实验中的可靠性，与真实人类数据对照，并揭示了失效条件，对关注仿真偏差和边界的研究者具有重要参考价值。","inspiration":"借鉴其多模型、多输入格式的因子设计，以及通过操纵检查和主题建模分离“看见”与“感知”的方法，可迁移到经济金融中的视觉信息处理场景，如央行沟通中的图表设计、金融产品广告或信贷审批中的视觉线索。｜设计一个实验，用 LLM 模拟投资者，呈现不同颜色或形状的金融图表（如涨跌颜色、风险提示图标），测量其风险感知和投资决策，并与真实投资者实验数据对照，检验 LLM 是否复现视觉线索对风险偏好的影响。"}},{"id":"2609.25760","version":1,"title":"The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance","zh_title":"模拟社会的局限：后训练与调查微调如何抹除跨文化方差","abstract":"Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \\textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\\%) keeps half the human spread overall ($\\dr = 0.50$) and 11\\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.","authors":["Rojin Ziaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25760","pdf_url":"https://arxiv.org/pdf/2609.25760","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","算法保真度","跨文化调查"],"reason":"直接评估LLM仿真人类调查回答的分布保真度，使用WVS真实数据对照，并揭示后训…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":3,"question":"LLM后训练与调查微调如何影响其模拟人类调查回答时的跨文化方差保留？","design":"使用WVS第7波12国10,000个受访者-问题对，构建包含人口统计和Inglehart-Welzel文化维度的价值编码persona，评估11个零样本LLM和5个在WVS数据上微调（SFT、DPO、GRPO）的变体，测量点准确率、MAE、Wasserstein-1距离和偏差比（预测标准差/人类标准差）。","baseline":"世界价值观调查（WVS）第7波12个国家的真实个体回答分布，包括标准差和分布形状。","findings":"后训练导致“共识坍缩”：监督指令微调使偏差比从1.22降至0.59，准确率仅提升0.9个百分点；后续DPO和GRPO未恢复方差，且WEIRD与非WEIRD国家差距扩大，最准确模型（Tulu 3 70B-DPO微调）整体偏差比0.50，尼日利亚仅0.11。提高采样温度至1.0不改变Wasserstein-1距离，GRPO在准确率或分布形状奖励下均不能恢复方差，混合未对齐先验仅将整体偏差比从0.51提升至0.62，尼日利亚仍为0.36。","reliability":"论文承认共识坍缩在非WEIRD国家更严重，且无法通过提高温度或GRPO恢复；混合未对齐先验只能部分缓解，不能根本解决。未讨论其他局限。","relevance":"直接针对LLM仿真人类调查回答的分布保真度，使用真实WVS数据对照，揭示后训练和微调对多样性的系统性压缩，对关注仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其诊断框架：同时测量点准确率和分布离散度（偏差比、Wasserstein距离），并沿后训练轨迹分解方差损失来源。｜可迁移到经济政策评估中的异质性反应仿真，如不同文化背景下对税收或福利政策的态度分布。｜用LLM模拟多国受访者对政策的态度，以WVS或类似跨国调查为真实基准，比较零样本、指令微调和偏好优化模型在点准确率与方差保留上的权衡，并检验温度调节和先验混合能否恢复分布。"}},{"id":"2609.25059","version":1,"title":"Can Large Language Model-Generated Responses Support Assessment Development? A Human-Calibrated Rasch Benchmark","zh_title":"大语言模型生成的回答能否支持评估开发？一项人类校准的Rasch基准研究","abstract":"Large language models (LLMs) are proposed as synthetic respondents for pilot testing, but their usefulness depends on whether they supply the evidence assessment development requires. We calibrated rating scale models on 14 digital-use skill items from 6,245 adults and used the human item parameters to evaluate responses generated for 1,300 demographically matched personas. LLM responses had high internal consistency ($\\alpha \\approx .94$) but used the lowest category in 0.2-0.3% of responses versus 13.7-21.5% for humans, and no persona chose it on every item. These gaps changed the response-category judgment on both subscales and the PC targeting judgment; on the human-calibrated scales, 12 of 14 items had infit below 0.70, indicating responses more predictable than the Rasch model expects. Preregistered changes to the prompt, category order, and persona information did not restore the human lower range. High internal consistency is insufficient evidence that LLM responses can replace human pilot data.","authors":["Eunjeong Song","Sehee Hong"],"categories":["stat.AP","cs.CY"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25059","pdf_url":"https://arxiv.org/pdf/2609.25059","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真被试","Rasch模型","人类数据对照"],"reason":"用LLM生成合成被试回答，并与真实人类数据校准，评估其替代可行性，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":6,"question":"LLM生成的回答能否支持评估工具开发中关于反应类别结构、目标定位和项目审查的判断？","design":"使用LLM为1300个人口统计学匹配的虚拟人物生成对14个数字技能自评项目的回答，并改变提示指令、类别顺序和人物信息以检验对低类别使用的影响。","baseline":"来自2025年数字鸿沟调查的6245名成年人的人类回答，用于校准Rasch模型并作为固定参照。","findings":"LLM回答内部一致性高（α≈.94），但最低类别使用率远低于人类（0.2-0.3% vs 13.7-21.5%），且没有虚拟人物在所有项目上选择最低类别。在人类校准的量尺上，14个项目中有12个的infit低于0.70，表明LLM回答比Rasch模型预期的更可预测，且提示、类别顺序和人物信息的改变未能恢复人类低端分布。","reliability":"论文指出高内部一致性不足以证明LLM回答可替代人类试点数据；人口统计学匹配不能保证再现测量构念的变异；研究为回顾性基准，不适用于儿童或青少年。","relevance":"该研究直接针对LLM作为合成被试的可靠性，用真实人类数据校准并发现系统性偏差，对关注仿真失效条件的研究者具有重要参考价值。","inspiration":"借鉴其用人类校准的Rasch模型作为固定参照来评估LLM回答偏差的方法，可迁移到经济金融中的主观量表或调查数据仿真（如消费者信心、风险态度）。｜可应用于政策评估中的问卷预测试，例如用LLM模拟不同人口群体对政策的态度分布。｜设计：以真实家庭调查（如美国消费者财务调查）为基准，用LLM生成匹配人口特征的虚拟受访者对风险偏好或通胀预期问题的回答，比较分布差异和项目反应模型拟合，检验提示工程能否校正偏差。"}},{"id":"2609.24012","version":2,"title":"Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure","zh_title":"检验而非假定充分性：针对涌现网络结构校准生成式社会模拟器","abstract":"Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.","authors":["Tengfei Shao","Chao Li","Xu Wang","Masayuki Goto"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-22","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.24012","pdf_url":"https://arxiv.org/pdf/2609.24012","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","网络校准","市场模拟"],"reason":"用LLM生成persona模拟市场网络，并与真实数据校准，评估仿真充分性，属核…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":7,"question":"如何对生成式社会模拟器进行校准与充分性检验，使其能可靠地复现真实市场的涌现网络结构？","design":"使用基于LLM一次性离线生成的人物画像构建前向模型，模拟二手奢侈品转售市场中买家与品牌的二分网络；通过摊销后验估计校准行为参数，并进行合成可识别性评估、匹配样本量的充分性检查、诊断引导修复和留出统计量审计。","baseline":"真实二手奢侈品转售市场的交易数据，按渠道和居住地划分为四个单元，每个单元为一个买家-品牌二分网络。","findings":"行为参数在所有四个单元中可恢复，但校准近似且对一个参数过度自信；所有单元中观测摘要均超出模拟器的可达性参考，平均购买层级是普遍差异，修复未能恢复充分性，留出审计发现买家广度离散度的遗漏。","reliability":"论文承认校准近似且对一个参数过度自信，修复未能恢复充分性，LLM画像的价值在于结构而非品牌身份，且不声称模拟器再现了市场。","relevance":"该研究直接针对LLM社会模拟的校准与充分性检验，提供了与真实数据对照的严格方法，对关注模拟可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其将模拟器校准与充分性检验结合的方法，通过留出统计量审计避免循环验证，并利用合成可识别性评估参数可恢复性。｜可迁移到消费者市场细分与品牌选择模拟，例如用LLM生成消费者画像模拟电商平台上的购买行为，并与真实交易数据对照。｜以LLM生成消费者画像作为被试，施加不同营销策略（如折扣、推荐）作为处理，结果变量为购买品牌网络结构，用真实电商交易数据作为对照，评估模拟器能否复现品牌集中度与社区结构。"}},{"id":"2609.25586","version":1,"title":"Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus","zh_title":"偏转价值罗盘：与大语言模型互动暂时将人类价值优先转向个人关注","abstract":"Large language models increasingly support decisions where values are in tension, yet little is known about whether interacting with them changes which values users prioritize. In a preregistered study, 200 U.S. adults interacted with ChatGPT, Claude, or Gemini as a thinking partner or read fixed AI-generated considerations. The prompt asked LLMs to support reasoning without recommending a decision and named no values. Participants advised people facing real dilemmas and completed parallel PVQ-RR forms before, immediately after, and one task later. Each LLM condition temporarily shifted value priorities toward personal focus relative to the control (d=0.37-0.51), primarily through increased Self-Enhancement. Participants' advice retained words and meaning from their exchanges. Thus, a brief LLM interaction that neither targets values nor seeks to persuade can reorient values active during judgment without detectable convergence in value directions or advice.","authors":["Hasibur Rahman","Malak Sadek","Smit Desai"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25586","pdf_url":"https://arxiv.org/pdf/2609.25586","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM影响人类","价值观转变","人机交互实验"],"reason":"研究LLM交互对人类价值观的影响，有真实人类对照，揭示仿真偏差条件","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":8,"question":"与LLM进行简短的思考伙伴式互动是否会暂时改变人们在判断中激活的价值优先级，以及这种改变是否持续、是否导致价值或建议的趋同？","design":"本研究不是用LLM模拟人类被试，而是以200名美国成年人为真实被试，随机分配到ChatGPT、Claude或Gemini三种LLM交互条件，或一个阅读固定AI生成考虑的非交互控制条件。在实验第二阶段，被试与LLM进行最多10分钟的思考伙伴式互动（提示词不提及任何价值观、不推荐决策），然后为真实困境提供建议并完成Schwartz的PVQ-RR价值观量表；在第三阶段，被试在没有LLM的情况下处理另一个困境，以测量效应的持续性。","baseline":"有真实人类对照：控制组被试阅读由三个LLM合成的固定AI生成考虑，但不与LLM交互；所有被试在互动前、互动后立即和后续任务中完成平行版本的PVQ-RR量表，作为价值观变化的基准。","findings":"与LLM互动后，被试的价值优先级相对于控制组暂时向个人关注方向偏移（效应量d=0.37-0.51），主要通过自我增强价值的增加实现；这种偏移在后续任务中消失，且未检测到价值方向或建议的趋同。","reliability":"论文承认效应是暂时的，在后续任务中消失；未讨论其他失效条件或局限。","relevance":"该研究直接探讨LLM交互对人类价值观的因果影响，属于批判性仿真研究，揭示了在无明确说服意图下LLM仍能暂时改变价值优先级，对理解LLM在决策支持中的潜在偏差具有重要意义，值得精读原文。","inspiration":"值得借鉴的做法是采用随机对照实验设计，将LLM作为处理条件，设置非交互控制组，并使用标准化的价值观量表在多个时间点测量效应。｜可以迁移到经济金融中的消费者跨期选择或投资决策场景，例如LLM作为财务顾问是否会影响个人的时间偏好或风险态度。｜一个可行的设计是：招募真实投资者作为被试，随机分配到与LLM（如ChatGPT）进行投资讨论的处理组或阅读固定建议的控制组，在互动前后测量时间贴现率和风险偏好（如使用滴定法或量表），并与真实市场数据（如实际投资组合选择）进行对照，以检验LLM交互对经济决策的因果影响。"}},{"id":"2609.23640","version":1,"title":"Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment","zh_title":"人类对齐模型是人类的模型吗？偏好对齐中的图灵测试差距","abstract":"Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.","authors":["Suqin Yuan","Runqi Lin","Muyang Li","Guanzhe Hong","Jindong Gu","Lei Feng","Chris Russell","Tongliang Liu"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23640","pdf_url":"https://arxiv.org/pdf/2609.23640","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好对齐","人类仿真","算法保真度"],"reason":"研究偏好对齐导致模型行为偏离人类，评估仿真可靠性，具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":9,"question":"人类偏好对齐是否会使语言模型的行为偏离人类行为分布，从而产生图灵测试差距？","design":"本研究并非直接进行人类仿真实验，而是通过理论分析和受控实验检验偏好对齐对模型人类相似度的影响。在受控实验中，使用同一批人类回答分别进行等权拟合和偏好加权拟合，比较模型对未见过的人类回答的似然；在DPO实验中，从已拟合人类回答的检查点出发进行偏好优化，测量人类回答似然和生成文本的变化。","baseline":"对照的真实人类数据来自SHP和StackExchange回答池，其中包含人类撰写的回答和人类投票偏好；以及HH-RLHF、WebGPT等偏好数据集。","findings":"偏好对齐在理论上仅在严格条件下保持人类行为分布，而实证中未发现人类偏好满足该条件。偏好加权强度越大，模型对人类回答的似然损失越大，且该损失与加权方向无关；标准DPO同样导致模型偏离人类行为。","reliability":"论文指出，提高人类相似度并非普遍可取，可能引发冒充、操纵和社会工程等风险；同时，通过角色提示和偏好人类回答的DPO等方法对人类相似度的恢复有限且依赖领域。","relevance":"该研究直接揭示了用偏好对齐的LLM作为人类被试替代品时可能存在的系统性偏差，为评估仿真可靠性提供了关键理论依据和实证证据，值得精读。","inspiration":"借鉴其将偏好对齐与行为分布分离的框架，在仿真实验中明确区分“人类偏好”与“人类行为”，并测量两者差异。｜可迁移到经济金融中的政策偏好调查或消费者决策仿真，例如用LLM模拟公众对通胀预期的回答，但模型可能给出“理想”而非“实际”的回答。｜设计：以LLM为被试，施加偏好对齐强度不同的处理，结果变量为对经济问题回答的分布与真实调查数据（如密歇根消费者调查）的KL散度，对照真实人类回答。"}},{"id":"2609.25572","version":1,"title":"A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators","zh_title":"行为特质泄漏到偏好中：诊断LLM用户模拟器中的特质干扰","abstract":"LLM-based user simulators aim to bridge the offline-online gap in recommender evaluation by emulating users through injected traits, where preference attributes determine what a user engages with and a behavioral activity trait governs how long they browse. However, we show this intended trait independence collapses during simulation, causing two failures: (i) Trait Interference, where amplified activity distorts preference boundaries and forces interactions with mismatched items to sustain browsing, and (ii) Evaluation Invalidity, where satisfaction scores inflate with activity-driven page counts despite taste mismatches, biasing evaluation toward trait distributions rather than recommender performance. To resolve this, we propose PQA, a page-level quality anchoring method that guides simulators using a personalized anchor reflecting each user's intrinsic preference standard. By assessing whether a page meets this standard before further browsing, PQA enables proactive exits from low-quality pages, letting the activity trait retain its intended role of modulating browsing depth within preference-conforming pages. Experiments show PQA mitigates trait interference and improves the reliability of LLM-based simulator evaluation under activity shifts. Our code is available at https://github.com/chaehyun1/PQA","authors":["Chaehyun Kim","Sein Kim","Hongseok Kang","Chanyoung Park"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25572","pdf_url":"https://arxiv.org/pdf/2609.25572","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM用户模拟","推荐系统评估","仿真偏差"],"reason":"用LLM模拟用户行为并诊断仿真失效，有真实数据对照，方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":11,"question":"LLM用户模拟器中，行为活动特质是否干扰了偏好特质，导致模拟失效？","design":"使用Agent4Rec和SimUSER两个LLM用户模拟器，在MovieLens和Amazon CDs数据集上，以SASRec为推荐模型生成20项推荐列表，按4项一页呈现。通过将低活动用户的活动特质修改为高活动状态，保持偏好特质不变，测量模拟器选择项目的语义对齐度、类别偏好一致性、浏览页数和满意度分数。","baseline":"真实用户数据：MovieLens和Amazon CDs数据集中用户活动水平与平均评分的负相关关系，即更活跃用户平均评分更低。","findings":"活动特质放大导致模拟器选择与用户偏好语义距离更远的项目，并偏离固有类别偏好，产生特质干扰；同时满意度分数随活动驱动的浏览页数增加而膨胀，即使推荐质量下降也持续浏览低质量页面，导致评估无效。","reliability":"论文指出当前模拟器缺乏页面级质量锚点，导致活动约束覆盖页面级偏好对齐；PQA方法依赖用户历史类别消费密度，可能受历史数据稀疏性影响，但未详细讨论其他失效条件。","relevance":"该研究直接诊断LLM模拟器中行为特质与偏好特质的干扰问题，并提供了真实人类数据对照，对关注仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文了解PQA方法细节。","inspiration":"借鉴其通过修改单一特质并保持其他特质不变来分离干扰效应的实验设计，可用于经济金融仿真中识别行为参数对决策的因果影响｜可迁移到消费者金融决策仿真，如信用卡使用或投资组合选择，检验风险偏好特质是否被交易频率等行为特质干扰｜设计：用LLM模拟消费者，将风险偏好设为固定，改变交易频率特质，测量投资组合风险水平和满意度，对照真实交易数据中频率与风险偏好的关系。"}},{"id":"2609.26403","version":1,"title":"AI-Generated Email Drafts Shift Culturally Distinctive Communication Styles in Professional Email","zh_title":"AI生成的邮件草稿改变职场邮件中文化特有的沟通风格","abstract":"AI assistants that support email composition may shift cultural communication norms, such as the directness typical of low-context cultures like the US versus the indirectness and contextual sensitivity central to high-context cultures like Japan. Yet it remains unknown to what extent people adopt and edit AI drafts inconsistent with their cultural communication norms. We address this through a preregistered within-subject experiment in which Japanese and American participants wrote workplace emails in their native language without AI, with a low-context AI, and with a high-context AI. We found that Japanese participants wrote emails with significantly more high-context markers (politeness, apologies) than Americans. But AI drafts shifted participants' emails toward the draft's style, with larger shifts when the draft was culturally misaligned: Japanese drifted most under low-context drafts, Americans most under high-context drafts. These findings suggest AI drafts risk overwriting cultural communication norms unless they adapt to users' communication styles.","authors":["Shintaro Sakai","Alice Gao","Yuichi Shoda","Katharina Reinecke"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26403","pdf_url":"https://arxiv.org/pdf/2609.26403","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","文化沟通","人机交互实验"],"reason":"用LLM生成邮件草稿并测量其对人类写作风格的影响，有真实人类实验对照，且揭示A…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":11,"question":"使用AI生成的邮件草稿是否会将用户的邮件写作风格转向草稿的文化语境风格，且这种转变在AI与用户文化背景不一致时是否更大？","design":"本研究不是用LLM模拟人类被试，而是用LLM生成邮件草稿作为实验干预。175名日本和美国全职员工在三种条件下写工作邮件：无AI辅助、低语境AI草稿、高语境AI草稿（被试内设计，条件顺序平衡）。结果变量是邮件中高语境标记（如礼貌用语、道歉）的数量。","baseline":"无AI辅助条件下参与者写的邮件作为人类基线，用于比较日本和美国参与者的文化沟通风格差异。","findings":"日本参与者在无AI条件下写的邮件比美国人包含更多高语境标记。AI草稿使参与者的邮件风格向草稿的文化语境偏移，且当草稿与参与者文化背景不一致时偏移更大：日本人在低语境草稿下偏移最大，美国人在高语境草稿下偏移最大。","reliability":"论文未讨论","relevance":"该研究用真实人类实验检验了LLM生成内容对人类行为的影响，有严格的人类对照，揭示了AI在文化维度上的潜在偏差和影响，对关注LLM仿真可靠性及偏差的研究者有参考价值。","inspiration":"借鉴其被试内设计和多条件对照，通过测量行为变化来评估AI干预的因果效应。｜可迁移到经济金融中的沟通与决策场景，如AI辅助撰写投资建议、信贷沟通或政策沟通，研究AI风格对用户决策或行为的影响。｜设计一个实验：招募金融从业者或普通投资者作为被试，让他们在无AI、低风险偏好AI草稿、高风险偏好AI草稿三种条件下撰写投资建议邮件，测量建议的风险水平，并与真实历史投资建议数据对照，检验AI是否改变建议风格及在风格不一致时影响是否更大。"}},{"id":"2609.21997","version":2,"title":"Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents","zh_title":"用于LLM智能体可控观点动力学的贝叶斯信念层","abstract":"LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior. We introduce Bayesian Chronicle Agents (BCA), a minimal belief layer separating what an agent believes from how it speaks. Each stance is a probability, updated by one Bayesian step per utterance heard. A single prior-strength parameter $\\kappa$ encodes stubbornness, modeled after its role in Friedkin--Johnsen (FJ) opinion dynamics. We then sweep this parameter to yield three canonical regimes of opinion dynamics on demand (consensus, persistent disagreement, committed-minority influence), with persistent disagreement matching the FJ closed-form fixed points at $R^2\\!=\\!0.93$--$0.99$. We further show that prescribed $\\kappa$ remains recoverable after the language round-trip, with perfect rank-order recovery across all four models. Explicit belief also makes simulation auditable: the layer surfaces systematic per-model stance biases that end-to-end simulation would silently absorb.","authors":["Hafsa Akbar","Daniel Platnick","Marjan Alirezaie","Hossein Rahnama","Alex 'Sandy' Pentland"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-09-21","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.21997","pdf_url":"https://arxiv.org/pdf/2609.21997","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["观点动力学","社会模拟","LLM智能体"],"reason":"用LLM agent模拟观点动态，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:35","error":null,"has_summary":false,"summary":null},{"id":"2609.26539","version":1,"title":"A retrospective analysis on the use of LLMs to study infant syntax learning","zh_title":"对使用大语言模型研究婴儿句法学习的回顾性分析","abstract":"Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.","authors":["H\\'elie Bazin (SCAI, SND, ISIR)","Anouk Barberousse (SND)","Fran\\c{c}ois Yvon (MLIA)"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26539","pdf_url":"https://arxiv.org/pdf/2609.26539","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM认知建模","婴儿语言习得","方法论评估"],"reason":"用LLM研究婴儿句法学习，属认知建模而非人类被试仿真，但涉及模型与人类对照，边…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.26579","version":1,"title":"Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models","zh_title":"接纳性而非谄媚：区分语言模型中的参与和顺从","abstract":"A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers of social sycophancy are also characteristic of conversational receptiveness, a construct from social psychology shown to improve interactions across disagreement. We argue that this overlap creates a construct-validity problem for social sycophancy evaluations. Using a popular moral-advice dataset, we find that responses classified as more socially sycophantic are also more receptive. Further, increasing the receptiveness of human-written responses---while preserving their substantive conclusions---causes them to be classified as more socially sycophantic. This tight coupling raises the possibility that social sycophancy evaluations inadvertently penalize desirable behavior. In a preregistered experiment comparing substantively equivalent responses, participants prefer the more receptive responses, expect users to be more likely to listen to them, and are more willing to seek advice from their authors. The same overall pattern persists even among participants who believe the original question asker is in the wrong. Finally, we introduce a simple approach that substantially increases receptiveness without increasing substantive deference, demonstrating that conversational receptiveness and substantive independence can be achieved together.","authors":["Calvin Isley","Johann Gaebler","Max Lamparth","Julia Minson","Sharad Goel"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26579","pdf_url":"https://arxiv.org/pdf/2609.26579","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为评估","社会心理学","模型对齐"],"reason":"研究LLM的社交行为与人类偏好，但非仿真人类被试，而是测量模型本身特性。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.26481","version":1,"title":"Behavior is Not Enough: A Mechanism-Based Evaluation of Social Norm Emergence in LLM Societies","zh_title":"行为不足：基于机制的LLM社会规范涌现评估","abstract":"Social norms cannot be identified from behavior alone: the same cooperative equilibrium may reflect shared expectations, strategic incentives, or simple imitation. Yet in multi-agent large language model systems, prior work largely treats behavioral convergence as evidence of norm emergence. In this work, we introduce an evaluation framework that measures agents' reported empirical and normative expectations in addition to behavioral convergence. Through controlled ablations, we test the effect of expectation elicitation and isolate two collective mechanisms central to theories of norm formation---social learning through interaction and social selection through network-based group formation. We further test the stability of these resulting dynamics under adversarial disruption across four LLM families. We find that eliciting expectations increases cooperative contributions, while social learning stabilizes behavior, and social selection reliably identifies cooperators but provides limited behavioral reinforcement. Following disruption, normative expectations and behavioral coordination recover differently. Together, these results show that similar cooperative outcomes can arise from different underlying social processes. By making expectations observable, our framework allows us to attribute each mechanism's contribution separately, offering designers of multi-agent systems a principled basis for selecting the social processes that sustain cooperation.","authors":["Rasika Muralidharan","Haewoon Kwak","Jisun An"],"categories":["cs.MA","cs.CL","cs.CY","cs.GT","cs.SI"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26481","pdf_url":"https://arxiv.org/pdf/2609.26481","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","规范涌现","多智能体"],"reason":"多智能体社会规范涌现模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.25194","version":1,"title":"Indirect tipping: a social attack surface in AI agent populations","zh_title":"间接引爆：AI智能体群体中的社会攻击面","abstract":"As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same equilibria that enable agents to coordinate also create a social attack surface. The standard framework to assess this vulnerability is critical mass dynamics: the minimum fraction of adversarial agents required to overturn an equilibrium through direct competition. Here, we show that this approach risks underestimating system vulnerability by reducing the problem to the identification of singular tipping points, and ignoring indirect but potentially more efficient routes through which collective behavior can be redirected. Through experiments with populations of LLM agents and an analytic framework that captures their collective dynamics at scale, we map critical-mass thresholds that define a directed, weighted topology over the space of coordination equilibria, and treat this topology as a navigable landscape. We show that indirect tipping through intermediate stepping-stone equilibria can reduce the committed minority required to reach an alternative state, bypass majority requirements, and make possible transitions inaccessible through direct challenges. The diversity of available alternatives and timing of the attack further reshape this landscape, creating opportunities for control as well as risks of unintended destabilization. These results show that an equilibrium's resistance to committed intervention is not an intrinsic property but a structural feature of its competitive relations with alternative states. Securing populations of interacting AI agents therefore requires mapping this social landscape alongside individual agent capabilities and the technical channels through which they interact.","authors":["Ariel Flint","Luca Maria Aiello","Sara M. Constantino","Romualdo Pastor-Satorras","Andrea Baronchelli"],"categories":["cs.MA","cs.AI","cs.CY","cs.SY","eess.SY","physics.soc-ph"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25194","pdf_url":"https://arxiv.org/pdf/2609.25194","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM多智能体","社会模拟","集体行为"],"reason":"用LLM agent群体模拟社会协调过程，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.25287","version":1,"title":"Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors","zh_title":"大语言模型能否识别和修复关系破裂？临床医生实践与LLM行为的比较","abstract":"Ruptures represent common albeit critical moments in interaction where relational alignment breaks down, making them essential for evaluating AI where trust and engagement matter most. In a scenario-driven empirical study, we examined the performance of three LLMs at identifying and resolving ruptures across 21 mental health conversations and 22 experts' evaluation of the strategies. For identification, LLMs relied on explicit linguistic cues within single turns whereas experts integrated implicit, relational, and contextual information across the conversation. For resolution, LLMs tended to produce more directive and scripted responses whereas experts adopted process-oriented strategies such as validation, open-ended exploration, and psychoeducation. Overall, LLMs showed higher agreement with predefined labels in identification, but not in resolution where experts rated their responses only moderately effective, with consistent limitations in timing, depth, and contextual sensitivity. We discuss implications for the design of mental health conversational agents emphasizing relational awareness, pacing, and human-in-the-loop support.","authors":["Jeongah Lee","Joy Qiuyue Zhong","Drishti Goel","Violeta J. Rodriguez","Dong Whi Yoo","Koustuv Saha","Ravi Karkar"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25287","pdf_url":"https://arxiv.org/pdf/2609.25287","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","心理健康对话","人机对比"],"reason":"研究LLM在心理治疗对话中的行为，与人类专家对照，但非仿真人类被试，而是评估模…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.25432","version":1,"title":"Tipping Points in LLM-Based Multi-Agent Systems: Stance on Climate Change Action","zh_title":"基于LLM的多智能体系统中的临界点：对气候变化行动的态度","abstract":"Because significant action to counter global warming requires massive public support, it is important to understand the dynamics of public opinion on climate issues. Of special interest are social tipping points, as revealed by large-scale effects of small perturbations in individual behaviors. Agent-based models (ABM) are an effective computational tool for studying these matters, because they allow controlled and systematic exploration of the effects of interventions that may be infeasible in real-world social systems. Large language models (LLMs) have been used to endow model agents with the ability to communicate in natural language (rather than by exchanging predefined messages), as well as with personality (in the form of a narrative self and episodic memory). We leverage LLM-powered ABM to look for tipping points in the social dynamics of a micro-society in which some of the discussions are about climate change. Our agents' stance was defined by two variables: the strength of conviction about the urgency of climate action and the degree of trust in existing institutions. We quantified shifts in agents' \"beliefs\" by monitoring, across multiple rounds of conversations, (1) inter-agent distances in this two-dimensional stance space and (2) the patterns of discussion topics as modeled by Latent Dirichlet Allocation (LDA). Our findings to date suggest that significant abrupt changes in climate-change stance do occur in this simple model. We report a number of methodological lessons from this study, notably, the need to prevent LLM biases from interfering with the conversational dynamics and, more generally, to maintain agent personality and episodic memories of interactions in the face of such biases. Resolving these issues may allow for using ABM-derived insights in designing real-life interventions vis-a-vis climate change and other important societal challenges.","authors":["Astghik Altunyan","Shimon Edelman"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25432","pdf_url":"https://arxiv.org/pdf/2609.25432","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["A3","D3"],"tags":["LLM智能体","社会模拟","舆论动态"],"reason":"用LLM智能体模拟社会舆论动态，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2607.22006","version":2,"title":"Printed but not benchmarkable: most building-decarbonisation disclosure cannot be matched to the pathways that stranding regulation assumes","zh_title":"已印刷但不可基准化：多数建筑脱碳披露无法匹配搁浅监管假设的路径","abstract":"Cities are beginning to enforce carbon limits on existing buildings. Science-based decarbonisation pathways set those limits one asset type and one jurisdiction at a time. Owners, however, report for the whole firm. We measure what that mismatch costs on two sets of public corporate reports: a census of 502 reports from the 119 listed built-environment firms with a collected report inside a 2,246-firm panel (2003-2023), and 519 real-estate reports from 101 firms (2007-2024). BeDA, a multimodal language-model tool whose reliability we test first, read them. Running the pathway frameworks' own entry tests over published disclosure: 16.5% of census reports (43.7% of real-estate reports) print an operational carbon intensity per square metre; 6.6% (25.0%) can be matched to a pathway for their property type in a covered jurisdiction; and only 5.0% (16.4%) disclose the floor area they divided by. Of the failures at the pathway test, 82-84% follow from reports lumping the portfolio together and 16-18% from a missing curve in the pathway library. The obstacle is the reporting unit, not missing data. The rate is roughly twice as high for European as for US listings (65-71% versus 35% in listed real estate). We also show that a US portfolio's carbon verdict cannot be worked out from disclosure at all. Within one climate zone, the pathway's carbon limit varies by up to 2.79-fold with the electricity subregion, which no report names; its energy limit does not move. Extraction is checked against the source PDFs (97.3% of extracted intensities appear verbatim) and repeats on a second extractor (kappa = 0.97). Recall of the non-disclosing class was 95.1% in a blinded hand audit of 122 reports. The fix follows from the measurement: split intensity by asset type and jurisdiction, and report floor area.","authors":["Jingyi Xu","Minghui Cheng","Anchen Sun"],"categories":["cs.CY","econ.GN","q-fin.EC","stat.AP"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-23","first_seen":"2026-07-27","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2607.22006","pdf_url":"https://arxiv.org/pdf/2607.22006","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","气候披露","文本提取"],"reason":"用LLM提取披露信息，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.25669","version":1,"title":"From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits","zh_title":"从话语到网络：建模俚语在Subreddit中的采纳与扩散","abstract":"Adoption and diffusion of neologisms in online communities have received renewed attention in recent years. As internet slang terms such as APT, referring to a K-pop song, and phrases such as Canon Event meaning an embarrassing but pivotal event, go viral online, it becomes increasingly important to understand the mechanisms that contribute to their success. Prior studies have often explained slang diffusion either from the perspective of social interaction or from the linguistic properties of the slang itself, but rarely from both perspectives together. One major obstacle has been the high cost of annotating slang usage in large-scale online communication. Recent advances in large language models (LLMs), however, make it possible to use them as scalable annotators for such tasks. In this study, we first curate a human-annotated benchmark to evaluate LLM performance in detecting slang usage in real Reddit communication. We then leverage LLM-based annotations to model slang adoption and diffusion. Our results show that slang diffusers with higher bridging capital are associated with increased subsequent adoption, whereas diffusers with higher bonding capital are associated with reduced adoption. We also find that wider contextual usage of a slang term is associated with a longer time before new users officially adopt it. Together, these findings suggest that both social-network structure and linguistic context shape the diffusion of neologisms in online communities.","authors":["Xiaoning Wang","Ted Underwood","Zhewei Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25669","pdf_url":"https://arxiv.org/pdf/2609.25669","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","社会网络","语言扩散"],"reason":"用LLM做标注替代人工，非仿真人类被试，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.26527","version":1,"title":"A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation","zh_title":"一种符号学感知的框架用于评估自然语言生成中的保真度与覆盖率","abstract":"When two texts describe the same expression, standard metrics based on lexical overlap or whole-text similarity may fail to detect meaningful differences in how that expression is framed. We propose a framework to evaluate semiotic alignment between texts, where a semiotic profile encompasses both the contextual meaning and the discourse references made salient by a text. Our approach yields two scores, Semiotic Fidelity and Semiotic Coverage, estimating how much of one text's profile is supported by the other and how much of the other's profile it recovers. Experiments show that coverage is typically lower than fidelity, and that alignment between LLMs and human-curated data is highest at low sampling temperatures, while higher temperatures reduce this alignment.","authors":["Lorenzo Zangari","Davide Picca"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26527","pdf_url":"https://arxiv.org/pdf/2609.26527","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["语义对齐","评估框架","人类数据对照"],"reason":"评估LLM与人类文本的语义对齐，非仿真人类被试，但涉及人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.26204","version":1,"title":"WatchPoint: Executable User Feedback for Real-World Agentic Web Development","zh_title":"WatchPoint：面向真实世界智能体网页开发的可执行用户反馈","abstract":"When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge scoring, or natural-language corrections, but few interact with the live application the way a developer would. We introduce WatchPoint, a simulated-user system that mimics real developer behavior by generating and executing diagnostic scripts against the running application, producing structured observations that guide the coding model's retry. Unlike prior approaches that target single-file edits or evaluate using non-executable metrics, we operate on Web-Bench, a benchmark of 50 multi-file web projects comprising 1,000 sequentially dependent tasks, verified by deterministic end-to-end tests. WatchPoint recovers 57.6% of the tasks it diagnoses, and a controlled user study confirms the simulation's realism: human testers achieve a comparable recovery rate (54.5%), providing evidence that automated diagnostic scripts can substitute for interactive human testing on sequential web development tasks. We further identify a pattern of capability gaps that governs when simulated-user feedback is helpful and when it should be withheld.","authors":["Guanqun Yang","Wei Yang","Xueqing Liu"],"categories":["cs.SE","cs.AI","cs.CL"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26204","pdf_url":"https://arxiv.org/pdf/2609.26204","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["智能体反馈","模拟用户","软件工程"],"reason":"用模拟用户替代人工测试，属于替代人类劳动而非仿真被试，但涉及人类对照，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.26069","version":1,"title":"RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty","zh_title":"RankCert：模拟学习者何时能安全选择AI导师？结构不确定性下的鲁棒决策认证","abstract":"Simulation-based tutor selection can be unstable when predictively adequate learner models imply different policy rankings. RankCert certifies one of eight equal-budget tutoring policies only when model-averaged utility, probability-best, posterior regret, cross-domain rank, family coverage, and leave-one-domain-out and leave-one-visible-family-out averages support the same candidate; otherwise it abstains. We evaluated RankCert in 1,280 frozen held-out settings spanning five rotating held-out oracle families, 64 scenarios per family, and four cohort sizes. Calibration used a licensed, de-identified EdNet-KT1 derivative with 5,000 learners and 590,056 retained responses; all five family representatives passed the frozen adequacy gate. Minimum-domain mean pairwise top-1 agreement was 0.272917 (95% CI [0.253646, 0.293229]), showing substantial structural disagreement. Cohort-noise variance decreased from n = 30 to n = 300, while the structural family share remained nonzero. RankCert reduced total held-out decision loss relative to full-coverage point selection by 0.006605 normalized-outcome units (95% CI [0.004859, 0.008407]). At comparable coverage, however, it did not reduce selective risk relative to a confidence-gated point certificate (difference -0.000213; 95% CI [-0.003238, 0.002384]; Holm p = 0.929654). Certification occurred in 3.75% of settings and only in stable scenarios; RankCert abstained in every ambiguous, misspecified, and structural-conflict setting. \"Safe\" denotes only benchmark-scoped decision certification under the declared utility and uncertainty set; no human-learning, causal, deployment-effectiveness, or general-safety claim is made.","authors":["Nizam Kadir"],"categories":["cs.AI","cs.CY","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26069","pdf_url":"https://arxiv.org/pdf/2609.26069","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["模拟学习者","AI导师选择","决策认证"],"reason":"用模拟学习者选择AI导师，属社会模拟但无真实人类行为对照，且非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.25790","version":1,"title":"Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence","zh_title":"用大语言模型自动化建构性评估：迈向可扩展和可重复的实践能力评价","abstract":"This study aimed to automate hierarchical diagnostic reasoning (HDR), a constructive method for evaluating practical judgment skills, by developing and testing an evaluation process using a large language model. HDR is a descriptive task that measures higher-order cognitive skills by requiring students to identify and explain errors in case-study-based problems. However, it requires expertise and effort to develop and evaluate. Hence, we proposed and empirically validated the automatic (1) generation of case problems containing errors aligned with educational intentions, (2) scoring of descriptive answers, and (3) generation of structured feedback based on incorrect answers, achieved solely through prompt design without fine-tuning. The internal consistency and construct validity of the generated problems were supported by the experimental score distribution and Cronbach's alpha (0.78). The agreement between automated and human ratings reached 100% under some conditions. The feedback was rated as being as convincing and useful as that from human instructors, demonstrating a practical framework for implementing HDR-based constructive assessment with reproducibility, immediacy, and low cost. The flexibility of large language models will also enable repeated and longitudinal assessments while maintaining structure, showing broad potential for application in educational settings.","authors":["Satoshi Takahashi","Atsushi Yoshikawa","Megumi Kose","Kenichi Suzuki","Chieko Inoue","Yumi Watanabe","Mari Sawada"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25790","pdf_url":"https://arxiv.org/pdf/2609.25790","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM自动评分","教育评估","建构性评价"],"reason":"LLM替代人工评分与反馈，属标注替代而非仿真人类被试，但涉及教育评估，需人工判…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2609.26707","version":1,"title":"Optimal Sequential Annotations for Off-Policy Evaluation","zh_title":"离线策略评估的最优顺序标注","abstract":"Offline reinforcement learning and off-policy evaluation evaluates dynamic treatment rules based on retrospectively collected data prior to deployment. In recent AI applications, state and reward information is recorded as complex text or image, which recent AI advancements such as LLM-as-a-judge can label with unknown bias. Expert annotation may be available but at a higher cost. For example, safety classification via cheap but imperfect classifiers vs. expensive expert review. We show how a limited budget for ground-truth data-annotation can be used via doubly-robust OPE with missing rewards, and we optimize variance-optimal annotation probabilities for sequential off-policy evaluation, where the target policy value is estimated from annotated data. We characterize the optimal annotation probabilities for sequential forward-monotone annotation protocols, and provide a feasible batch-adaptive implementation. Our work is motivated by a collaboration with a homelessness services nonprofit that writes casenotes for individuals over time. Our method can be used to unlock trustworthy inference from casenote data and answer new inferential questions such as: how does expanding outreach effort over time affect progress towards a housing application and improvement in housing placement? In simulations and on two real datasets - casenotes from the nonprofit and human-preference votes from LMArena - we see reductions in RMSE of 34-65% for housing placement and 17-68% for progress towards a housing application at budgets of 40% of full annotation and above, and by 55-62% at every budget on LMArena.","authors":["Woojin Chae","Ezinne Nwankwo","Haitong Qin","Angela Zhou"],"categories":["stat.ME","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26707","pdf_url":"https://arxiv.org/pdf/2609.26707","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["离线策略评估","LLM标注","主动学习"],"reason":"用LLM标注替代人工标注，属标注员替代而非仿真被试，但涉及人类数据对照与政策评…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":false,"summary":null},{"id":"2609.25284","version":1,"title":"When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning","zh_title":"当LLM智能体读不懂氛围：面向关系社会推理的ReAdapt","abstract":"A social agent's most basic decisions (should I react to this post? who should I reach out to?) are not purely content problems. The right action often hinges on the latent relationship between people -- tie strength, reciprocity, mutual connections -- rather than on which content is most salient. Standard LLM agent loops do not explicitly represent how new relational evidence should revise the agent's current social hypothesis, leaving them prone to surface-obvious choices when relational and content cues diverge. We formalize this failure mode with a relationship-reasoning benchmark: 500 synthetic social worlds with friendships, follows, reaction histories, and feeds, yielding 1,000 queries over two tasks, reaction selection and warm introduction (finding the best bridge to a target person). By construction, the surface-obvious candidate differs from the relationship-grounded oracle in about 53% of queries, forming an overturn subset where the agent must use relational evidence to revise an initially plausible choice. We propose ReAdapt (Relationship-Adaptive Agent with Policy-driven sTate), which augments the ReAct loop with an explicit structured social state z = (G, B, R, N, D) capturing goal, belief, relationship, norm, and disclosure. After each tool observation, ReAdapt runs a typed Adapt step that updates this state and emits a policy operation (continue, switch, abandon, or clarify) before choosing the next action. With Gemini-3-Flash on a stratified subset of n = 150 queries per task, ReAdapt improves warm-introduction accuracy from 37% to 51% (+14 points) and reaction-selection accuracy from 69% to 77% (+8 points). Oracle regret drops from 0.260 to 0.152 and from 0.095 to 0.053, respectively. Holding the model, tools, and environments fixed, these results suggest that explicit relational-state adaptation helps LLM agents turn retrieved social evidence into revised decisions.","authors":["Jianzhe Lin","Xiaolin Li","Yunda Liu","Fei Wang","Jubin Chheda"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25284","pdf_url":"https://arxiv.org/pdf/2609.25284","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","社会推理","关系适应"],"reason":"多智能体社会推理，无人类数据对照，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2608.27855","version":2,"title":"AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not","zh_title":"AI写作有稳定的风格指纹，但AI编辑没有","abstract":"Text generated by large language models (LLMs) has been shown to be stylometrically distinct from human-written text (Andre et al., 2023; Shah et al., 2023; Opara, 2024; Soto et al., 2024; Li and Zhang, 2025; Selvioglu et al., 2025). But LLMs are increasingly used not only to generate text but also to edit human writing, and it is unclear whether the two leave the same trace. We show that AI generation leaves a consistent \"stylometric footprint\": a small subset of features, primarily entropy and lexical diversity, consistently separates AI-generated text from human writing across 8 LLMs and 5 domains, while the remaining features depend heavily on the domain and generator. AI editing, however, does not reproduce the same footprint. Relative to their human- written sources, AI-edited texts show only a small increase in lexical diversity and a decrease in entropy, rather than the joint increase that characterizes AI generation. Lexical density, which contributes little to generation, instead becomes the dominant editing-associated signal. Stylometric features therefore separate AI-edited text from AI-generated text but are substantially less effective at separating it from human-written text. Our results suggest that \"AI text\" is not a single phenomenon: generation and editing leave qualitatively different stylometric traces and should be studied separately.","authors":["Zhengyang Shan","Yukyung Lee","Sophie Hao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-31","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.27855","pdf_url":"https://arxiv.org/pdf/2608.27855","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI文本检测","风格计量","文本生成"],"reason":"研究AI文本风格特征，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.08016","version":2,"title":"What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality","zh_title":"多智能体大语言模型辩论究竟改变了什么？分歧与答案质量的分层分析","abstract":"Multi-agent debate, in which several LLMs exchange arguments before producing an answer, is widely assumed to improve answer quality by surfacing genuine disagreement. That disagreement is hard to verify, and no single signal can settle it, so we organize the analysis around four questions: (A) does the debater say it disagrees; (B) does its reply text actually argue; (C) does the dissent survive once the tone instruction that produced it is removed; and (D) do the probabilities assigned to stance options change? We evaluate three-model committees on 50 open-ended GlobalOpinionQA questions under three debate tones, friendly (seek common ground), neutral, and hostile (stress-test every position). The answers differ. (A) Full agreement differs by 50.4 percentage points between the friendly and hostile endpoints, pooling replies across all three rounds. (B) The text argues too: an LLM evaluator reading the contribution and reply without the structured self-report or tone instruction confirms the pushback is real. (C) Removing the instruction produces more returns to agreement in our sample, but the primary question-level test is inconclusive. (D) In a separate open-weight study, adjusted movement toward the opposing side is detected in two of three models, relative to a filler-based reference. This measures contextual response probabilities, not lasting belief change. For final answers, debate brings no measurable quality gain: an evaluator that judges each pair in both answer orders scores the debated answer no better than the same committee's no-debate answer on all 299 pairs, a test that detects severe but misses moderate damage in our checks. Taken together, LLM debate readily changes what agents say, but we find much weaker evidence that it changes what they persistently endorse or improves the quality of the final answer.","authors":["Chen Qian"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-09","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.08016","pdf_url":"https://arxiv.org/pdf/2609.08016","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体辩论","答案质量","分歧分析"],"reason":"研究多智能体辩论对答案质量的影响，不涉及人类行为仿真或对照，属纯多智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":false,"summary":null},{"id":"2609.26035","version":1,"title":"Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance","zh_title":"可信AI的真相：表达怀疑、来源和信念修正作为可工程化立场","abstract":"Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and explicit belief revision can be implemented as a behavior layer over a fixed language model; we do not test believability or trust. The layer combines three epistemic states, per-claim confidence and typed provenance, a provenance-gated expression rule, and a persistent revision store with auditable acknowledgments and partial resistance to false corrections. We evaluate it on a constructed, mechanically scored multi-session benchmark using a synthetic model and Qwen2.5-0.5B-Instruct. The synthetic instrument passes all five checks. On the real model, acknowledgment soundness, a by-construction guarantee, holds in 100% of cases, and true corrections are accepted more often than false ones (0.44 vs. 0.15 on held beliefs; 0.875 vs. 0.420 including rule-accepted corrections of unheld facts), but the pre-specified expression-fidelity, contradiction-separation, and provenance margins fail. A disclosed post hoc analysis shows that expression gated on mean answer-token probability ranks correctness below chance end to end (AUC 0.41, conversation-clustered), whereas gating on sampling consistency discriminates (AUC 0.66). A consistency-gated configuration selected from this finding and evaluated under a separately committed protocol meets the conversation-level manipulation and capability-equivalence criteria and replicates on a redrawn conversation set. The manipulation result is selection-dependent, and both criteria remain unresolved when uncertainty is clustered over the 60 facts. The supported conclusions are limited to the by-construction audit guarantee, store-dependent partial correction discrimination, and a benchmark- and model-specific failure of token-probability gating; scaling the fact base is required before human evaluation.","authors":["Sebastian Cochinescu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26035","pdf_url":"https://arxiv.org/pdf/2609.26035","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话代理","可信度","信念修正"],"reason":"研究对话代理的可信度表达，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.26090","version":1,"title":"SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children","zh_title":"SpecialEduBench：面向自闭症儿童语言干预的视觉语言模型知识、技能与态度基准","abstract":"Language is the target of most early intervention for autistic children. Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds. Artificial intelligence is now being brought to that work, yet the benchmarks that reach special education ask what a model knows rather than what it does in front of a child. Building one is not straightforward, since whether a response is good teaching depends on what the child has just done, so no answer key applies. The evidence that settles it is visual as much as verbal, since the length of a wait, a shift of gaze, and the child's uptake leave no trace in a transcript. We introduce \\emph{SpecialEduBench}, which measures pedagogical competence along knowledge, skill, and attitude, with 4,537 knowledge items and with 200 skill items and 68 attitude items built on recorded intervention, the attitude items crossing pressure with monitoring into 192 response cells. Seven special-education experts wrote, scored, and reviewed the items, and we revised the judge model's instruction against the reference scores they set. Across eight frontier vision-language models no axis is saturated, since the strongest still fails about a tenth of the honesty cells. The models converge where the knowledge is factual and separate where the task is situated, and the failures gather where pressure is applied. We intend the benchmark as an audit to run before deployment and as a starting point for models built for this domain.","authors":["Jihoi Na","Taeyeong Kim","Sungjune Kong","Jaemin Jung","Min Joung Park","Kyungtae Joo","Ahhyun Kim","Shim Jaechang","Sooyoung Joo","Dongjin Ka","SeJoong Kim","Jimin Kim","HyunJin Jung","Unggi Lee"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26090","pdf_url":"https://arxiv.org/pdf/2609.26090","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准评测","视觉语言模型","特殊教育"],"reason":"纯模型能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.26210","version":1,"title":"Same Chart, Different Story: Bias in Vision-Language Chart Interpretation","zh_title":"同一图表，不同叙事：视觉语言模型图表解释中的偏见","abstract":"Vision-language models (VLMs) are increasingly used to interpret charts and generate natural-language explanations for socially consequential data. However, they may produce different narratives for the same chart when only the referenced social group changes, reinforcing stereotypes and misleading decisions. Despite these risks, no benchmark exists for systematically evaluating bias in chart interpretation across social dimensions. We introduce ChartBias, the first benchmark for auditing bias in VLM-based chart interpretation. ChartBias contains 820 manually curated real-world charts spanning six attributes: race, income, age, religion, immigration status, and gender, yielding 4,319 valid chart, attribute instances and 8,638 paired generations where the chart is fixed and only the group term is swapped. Across 12 proprietary and open-source VLMs, totaling 155,484 model responses, we find three widespread failure modes: narrative shift (same chart, different narratives), group hallucination (assigning a chart to a group without evidence), and preference polarity (favourable trends often linked to one group). We further propose a multi-agent mitigation framework that serves as a strong baseline by separating chart-grounded evidence extraction from group-conditioned generation and using a counterfactual judge to verify that group-driven differences are supported by the chart. The framework substantially reduces narrative shift while preserving chart-grounded reasoning. Our findings show that evaluating chart understanding requires measuring not only accuracy, but also fairness and consistency across social groups. We release ChartBias at https://github.com/vis-nlp/ChartBiasBench.","authors":["Mizanur Rahman","Huan Wu","Arash Asgari","Enamul Hoque Prince","Laleh Seyyed-Kalantari"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26210","pdf_url":"https://arxiv.org/pdf/2609.26210","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["VLM偏见","图表理解","公平性评测"],"reason":"评估VLM图表解释中的偏见，属于模型公平性评测，不涉及用LLM仿真人类被试或与…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:27","error":null,"has_summary":false,"summary":null},{"id":"2609.26399","version":1,"title":"Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding","zh_title":"结合层次化认知过程与过程监督的可解释场景安全理解","abstract":"Scene safety understanding plays a life-or-death role in situational awareness in various critical domains. Traditional methods that rely on learning direct mappings between scenes and safety levels often lack interpretability, limiting their reliability in critical applications. An effective approach to overcoming this challenge lies in interpreting human cognitive processes and equipping machine models with analogous cognitive capabilities. This work explores an effective way of integrating scene safety cognitive process modeling and process supervision. Specifically, we first construct a hierarchical cognitive safety structure, which motivates the development of a novel, high-quality scene safety understanding dataset based on multi-step reasoning with process labels. This dataset serves both as a benchmark and a resource to improve the safety reasoning capabilities of Large Language Models (LLMs), while also enabling a granular analysis of intermediate reasoning steps through information flow and saliency-based techniques. Building upon this foundation, we introduce a modular and flexible process supervision framework that reflects the hierarchical nature of human cognition. This framework leverages LLMs as the core architecture and incorporates Low-Rank Adaptation(LoRA) and Mixture-of-Experts (MoE) strategies to enable specialization and collaboration among expert modules, each tasked with specific sub-processes of the overall reasoning chain. Systematic experimental evaluations and analyses confirm that our framework exhibits superior interpretability and performance characteristics compared to traditional approaches.","authors":["Zhiyun Jiang","Hanyong Wang","Binbin Liang","Yu Xie","Zhengjie Wang","Menglong Yang","Wei Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26399","pdf_url":"https://arxiv.org/pdf/2609.26399","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["场景安全理解","可解释AI","过程监督"],"reason":"论文聚焦场景安全理解，用LLM做推理，但无人类被试仿真或行为对照，属自动驾驶/…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.26489","version":1,"title":"Calibration as a First-Class Criterion in LLM Evaluation","zh_title":"校准作为LLM评估中的首要标准","abstract":"Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.","authors":["Mario Sanz-Guerrero","Katharina von der Wense"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26489","pdf_url":"https://arxiv.org/pdf/2609.26489","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM校准","模型评估","可信度"],"reason":"论文讨论LLM校准评估，属NLP评测，不以人类行为为参照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.26629","version":1,"title":"PERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation","zh_title":"PersonaWeaver：程序化角色生成中超越传统原型的可控多样性","abstract":"Procedural character generation aims to populate games, simulations, and other virtual worlds with diverse characters. Large language models (LLMs) offer a promising foundation for scaling this task. However, LLM-based procedural character generation remains at an early stage: existing methods either generate characters directly or adapt profiles retrieved from persona banks. As we show, both approaches produce behaviorally homogeneous populations: characters overwhelmingly agree with positive moral norms and respond to questions with helpful, assistant-like reactions. To mitigate this homogenization, we introduce PersonaWeaver, which disentangles world building from behavioral specification and models behavior through setting general, diverse, manually curated banks of moral positions and conversational reactions. This design allows us to test how far LLM(s) can be pushed beyond their default behavioral patterns across settings. Across ten realistic and fantastical settings and three LLM(s), PersonaWeaver produces broader moral and interactional response distributions than prior work. Its guidance also diversifies interpersonal language, response length, and sentiment. It also produces less archetypal combinations of world attributes. Code is available at https://github.com/mqraitem/PersonaWeaver.","authors":["Maan Qraitem","Kate Saenko","Bryan A. Plummer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26629","pdf_url":"https://arxiv.org/pdf/2609.26629","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色生成","多样性控制","游戏NPC"],"reason":"生成游戏角色，无实验或测量目的，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:32","error":null,"has_summary":false,"summary":null},{"id":"2609.26687","version":1,"title":"Detecting GPT-Assisted Writing Using Interpretable Stylometric Features","zh_title":"使用可解释的文体特征检测GPT辅助写作","abstract":"Distinguishing GPT-assisted from independently authored student writing has become a critical challenge in academia. This paper evaluates the discriminative capability of interpretable stylometric features extracted solely from submitted text. Using data from 90 participants who wrote both independently and with ChatGPT assistance, we evaluate eight machine learning classifiers while keeping data from the same participant together during validation. On the held-out test set, Random Forest achieved an ROC-AUC of 0.87 and an F1-score of 0.84, with False Positive and False Negative rates of 22.2% and 11.1%, respectively. SHAP analysis shows that lexical and grammatical characteristics drive the resulting predictions. The findings suggest that transparent, text-intrinsic features provide measurable signal for detecting GPT-assisted writing.","authors":["Rajesh Kumar","Nabeel Siddiqui","Alexander Fuchsberger"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26687","pdf_url":"https://arxiv.org/pdf/2609.26687","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI写作检测","文体特征","学术诚信"],"reason":"检测AI辅助写作，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:33","error":null,"has_summary":false,"summary":null},{"id":"2609.25408","version":1,"title":"From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI","zh_title":"从离线代理到在线决策：面向对话式AI的分层参与度评估框架","abstract":"Online A/B experiments are the decision standard for user engagement, but traffic and readout time limit how many conversational-AI changes can be tested. We ask whether an offline signal designed to be computable without treatment-arm user exposure agrees with the outcomes of those experiments. We contribute a reusable construction and diagnosis checklist that treats an offline proxy as a chain of three alignments: behavioral label to product outcome, learned classifier to candidate-assistant behavior, and aggregated offline signal to experiment effect. A companion evaluation protocol audits the whole composite by interval-aware decision agreement, which compares offline and online confidence intervals instead of point estimates, and by within-experiment ranking. The instantiation we evaluate comprises a fixed evaluation suite on which candidate behavior is scored, an engagement classifier trained to predict session/prompt level engagements, and a calibration layer mapping sample-level score differences to online model-level engagement deltas. We then report the audit: 489 paired offline-online contrasts (one candidate arm against its control) from 27 experiments on a deployed multi-turn assistant, spanning model checkpoints to system-prompt tuning. Our primary test uses the 113 contrasts from eight experiments that ran after the map was frozen: on these the composite reaches 81.1% F1, against 34.3% for the raw classifier score it is built on, and makes no wrong-direction calls where that raw score makes 31. Every offline prediction was computed before its experiment ran to prevent overfitting. The evidence supports using the composite to prioritize candidates before scarce experiment traffic is allocated---in our deployment of the experiment, selecting among training checkpoints and tuning system prompts.","authors":["Xuanyi Li","Vaskar Nath","Hossein Amirkhani","Jay Li","Alex Deng"],"categories":["cs.AI","cs.IR","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25408","pdf_url":"https://arxiv.org/pdf/2609.25408","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["对话式AI","离线评估","A/B实验"],"reason":"研究离线代理与在线实验的一致性，属系统评估，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.26057","version":1,"title":"Observing the Conduct of Systematic Reviews with Generative AI Support: An Experience Report from a Graduate Software Engineering Course","zh_title":"观察生成式AI支持下的系统综述实施：一项研究生软件工程课程的经验报告","abstract":"Context: Secondary studies are fundamental practices in Evidence- Based Software Engineering, but teaching them requires activities that expose students to authentic methodological decisions. Objective: This paper reports an experience in a graduate course in which ten doctoral students in Software Engineering, organized into three groups, piloted secondary studies with and without support from generative AI. Method: A single-day classroom session was organized and observed, in which the groups conducted pilot systematic reviews with and without generative AI support. Classroom observations, produced artifacts, and interaction threads with assistants configured in ChatGPT were analyzed to reconstruct how each group appropriated the technology throughout the activity. Results: LLMs reduced initial barriers, accelerated the generation of alternatives, and made methodological problems more explicit, but they also favored excessive delegation, superficial validation, operational difficulties, and a shift in focus from conducting the SLR to using the tool. Conclusion: The experience offers a situated, observational account of how doctoral students engaged with generative AI during a systematic review activity, and the resulting insights also inform the design of a subsequent controlled study. The findings indicate that generative AI can support practical learning about SLRs, provided that its use is accompanied by human supervision, decision records, and critical reflection on its limitations.","authors":["Danilo Monteiro Ribeiro","Gilberto Sussumu Hida"],"categories":["cs.CY","cs.AI","cs.SE"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26057","pdf_url":"https://arxiv.org/pdf/2609.26057","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["生成式AI","系统综述","教学经验"],"reason":"研究LLM辅助系统综述教学，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.26562","version":1,"title":"The Disciplinary Language Transfer Problem: How Psychological Vocabulary Produces Governance Failures in AI Agent Deployment","zh_title":"学科语言迁移问题：心理学术语如何导致AI代理部署中的治理失败","abstract":"The vocabulary used to describe AI agents in governance contexts -- learning, memory, values, compliance, identity, trust -- is borrowed from psychological and organizational science, contributing to systematic failures in how organizations deploy, oversee, and hold agents accountable. This paper argues that the problem is not merely terminological but epistemological: psychological vocabulary carries an \"invisible grammar\" of its home discipline into governance discourse, calibrating frameworks to a metaphysical entity that does not exist in current AI architectures. We call this the disciplinary language transfer problem. Drawing on Wittgenstein's concept of language games, Kuhn's paradigm-laden observation, Haraway's situated knowledge, and Star and Griesemer's boundary object theory, we show that the transfer operates at three levels (epistemological assumptions, theoretical constructs, and surface vocabulary), each requiring a different remediation. We characterize six foundational epistemological assumptions embedded in Western psychological governance discourse, trace their origin in specific philosophical traditions, and show why each fails when applied to systems without developmental continuity. The paper's practical output is an actionable Disciplinary Audit: a six-question governance document scan operationalized through a translation taxonomy of thirty-seven terms mapping operational constructs to agent-appropriate replacements, presented here in abridged form and openly archived in full. The vocabulary reform proposed here is not merely terminological; it is the condition of possibility for governance frameworks that correctly identify what they are governing.","authors":["Kymberly Lasser-Chere","Tyler Akidau","Marc Millstone"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26562","pdf_url":"https://arxiv.org/pdf/2609.26562","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","语言哲学","概念分析"],"reason":"讨论AI治理中的语言问题，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.25311","version":1,"title":"\"I Talked an AI Chatbot, So What's Next?\" How U.S. Young Adults Imagine Responsible AI for Emotion Coping","zh_title":"“我和AI聊天机器人聊过了，接下来呢？”美国年轻人如何想象负责任的情绪应对AI","abstract":"Emotion coping is inherently relational, unfolding through interactions with friends, family, professionals, and communities. Yet AI chatbots are largely designed around a user--AI dyad. Learning from the ethics of care, we examine how AI chatbots shape the relational conditions of emotion coping. We conducted a scenario-based study with 17 U.S. young adults across four emotion coping scenarios. Participants identified eight roles through which AI could support relational conditions, alongside three challenges: flattening distinct relational conditions, discouraging reciprocity, and shifting relational labor onto users. This study contributed a relational perspective of responsible AI in emotion coping. We argue that responsible AI should respond to individuals' situated relational conditions rather than provide general-purpose support. We further identify two design principles: fostering reciprocity by supplying materials to engage with others, and strengthening emotional self-efficacy. Together, we position responsible AI as AI in the loop of human relationships.","authors":["Jiaying Liu","Nimra Ishfaq"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25311","pdf_url":"https://arxiv.org/pdf/2609.25311","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情绪支持","负责任AI"],"reason":"研究AI聊天机器人用于情绪应对，属于角色扮演对话，无实验或测量目的，不涉及LL…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.25700","version":1,"title":"Harnessing LLMs Without Surrendering Control: Delegation Boundaries in Visual Data Storytelling Authoring","zh_title":"在不放弃控制权的前提下利用大语言模型：视觉数据叙事创作中的委托边界","abstract":"Despite the emergence of large language models (LLMs) for visual data storytelling workflows, there are open questions about how authors decide what activities or tasks to entrust to them and what should be \"protected\" or maintained under human control. To investigate this, we interviewed a cohort of 12 expert visual data storytellers. Our analysis shows that participants rarely treated LLMs as autonomous storytellers. Instead, they tend to selectively delegate execution-oriented tasks to LLMs while retaining control over activities that shape narrative intent and story meaning. Our findings show that LLM assistance is most productive after human seeding and constraint-setting, and that it shifts labor from production to verification. We discuss design implications for boundary-aware authoring tools, data-grounded generation, low-fidelity ideation, and reporting practices for LLM-based visualization research. Supplemental materials for this paper are available at https://osf.io/hcnp6.","authors":["Zhuojun Jiang","Yuki Ueno","Chris Bryan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25700","pdf_url":"https://arxiv.org/pdf/2609.25700","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机协作","可视化叙事","LLM辅助创作"],"reason":"研究人类作者如何委托LLM执行可视化叙事任务，属于人机协作设计，非用LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.25774","version":1,"title":"Scientific capabilities and deployment sustainability of small-scale LLMs in biological wastewater treatment","zh_title":"生物污水处理中小规模大语言模型的科学能力与部署可持续性","abstract":"Large language models (LLMs) are emerging as scientific assistants, yet their computational demands and limited domain specialization constrain sustainable deployment in environmental engineering. Here, we investigate whether domain-specialized small-scale LLMs can combine scientific capability with sustainable deployment in biological wastewater treatment. We developed a benchmark evaluating three scientific capabilities of LLMs: retrospective cognition, comprehension fidelity, and prospective extrapolation. BioWater (8 billion parameters, fine-tuned on specialized domain knowledge) achieved higher comprehension-fidelity scores than participating human experts and performance comparable to a 397-billion-parameter general-purpose LLM in retrospective cognition and prospective extrapolation. Human-BioWater collaboration generated a scientific hypothesis that was subsequently supported by laboratory experiments, demonstrating its potential to contribute to prospective scientific research. We further evaluated the economic and environmental implications of LLM deployment across global wastewater treatment plants (WWTPs). Locally deployed small-scale LLMs became more sustainable than cloud-based large-scale LLMs as inference demand increased in intelligent WWTPs. These findings highlight domain-specialized small-scale LLMs as a promising pathway towards scientifically capable, computationally efficient, and sustainably deployable artificial intelligence for wastewater treatment.","authors":["Run-Ze Xu","Chu-Kuan Jiang","Dylan Ming-Han Li","Hong-Xiao Guo","Jia-Shun Cao","Guang-Hao Chen"],"categories":["cs.CE","cs.CY","stat.AP"],"primary_category":"cs.CE","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25774","pdf_url":"https://arxiv.org/pdf/2609.25774","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM科学助手","污水处理","领域专业化"],"reason":"论文聚焦LLM作为科学助手解决工程问题，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.25569","version":1,"title":"SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling","zh_title":"SambaGraph：用于足球战术响应建模的动作-反应时空图","abstract":"Soccer tactics are interactive: an attacking action changes the opponent's defensive problem, and the observed response depends on the multi-agent match state. We introduce SambaGraph, an action--reaction spatio-temporal graph dataset and benchmark for soccer tactical response modeling. From tracking and event data for all 64 matches of the 2022 FIFA World Cup, we curate 4,070 action-centered episodes represented as temporally aligned 23-node player--ball graph sequences with attack/defense views, response labels, and 26,270 split-safe attack--defense pairs. We study three questions: whether observed responses can be classified from graph episodes, whether successful defenses can be retrieved for a query attack, and whether graph-derived summaries support grounded LLM reasoning. A compact signature MLP obtains $0.796\\pm0.007$ macro-F1 for response classification, while a fused graph--signature dual encoder reaches $0.471\\pm0.029$ Hit@5 and $0.655\\pm0.051$ Hit@10 for full-bank defensive retrieval. Hard negatives maximize pair discrimination but not retrieval quality. Local LLMs underperform supervised encoders for direct classification and do not improve over a strong original order in eight-candidate reranking, but they provide grounded tactical rationales. These results position SambaGraph as a reproducible benchmark for graph-based soccer strategy-response research. Code and dataset are available at: https://github.com/areyesan/SambaGraph.","authors":["Abel A. Reyes-Angulo","Henry O. Velesaca","Steven Araujo"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25569","pdf_url":"https://arxiv.org/pdf/2609.25569","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["足球战术","图神经网络","LLM推理"],"reason":"足球战术建模，非人类被试仿真，LLM仅用于推理辅助","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2608.16886","version":2,"title":"Evaluating Beyond the Screen: Collective Assessment of AI-Generated Business Plans with Resource-Constrained Entrepreneurs","zh_title":"超越屏幕的评估：资源受限创业者对AI生成商业计划的集体评估","abstract":"Entrepreneurs increasingly use end-user generative AI technologies such as ChatGPT for high-stakes documents like loan applications and business plans, where AI-generated errors---a wrong price, a fabricated product---can affect loan or funding outcomes. Current approaches to supporting evaluation of AI-generated text assume a single user assessing output alone, on screen. This can be especially demanding for resource-constrained entrepreneurs, whose digital and AI skills vary widely. In this early-stage work, we explore how evaluation might instead be organized in a group setting and completed as a collective activity. We extended BizChat, an AI-powered business-planning tool, with an evaluation module that links each generated claim to the entrepreneur's original input. We partner with community organizations in Maryland---embedding BizChat within various entrepreneurship programs---where workshop attendees (N=14) evaluated their plans through think-pair-share discussion. Early findings suggest interface scaffolds like claim-to-input links primed attendees with concrete, personal evaluations, which the group setting then extended beyond the screen: attendees requested printed copies, used rubrics to compare across plans, and drew on peers' knowledge to verify what they could not easily judge alone.","authors":["Qi Zhao","Marjory Pineda","Ketul Chhaya","Aakash Gautam","Yasmine Kotturi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-18","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.16886","pdf_url":"https://arxiv.org/pdf/2608.16886","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","AI评估","创业支持"],"reason":"研究人类如何评估AI生成内容，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":false,"summary":null},{"id":"2608.25581","version":2,"title":"Are Concept Bottleneck Models Effective as Decision-Support Systems?","zh_title":"概念瓶颈模型作为决策支持系统是否有效？","abstract":"Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,959) evaluating how concept-based explanations and user interventions on the model's concepts affect the performance of the human-AI team in two distinct binary classification tasks. Our results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance with non-interpretable AI support. However, these benefits emerge only under certain conditions: classification tasks perceived as difficult, easily identifiable concepts, and active interaction with the model. We also discuss how inaccurate concept detection may undermine users' trust in the model. Overall, this work provides practical guidance for the deployment of CBMs as effective decision-support tools.","authors":["Alessandro Bogani","Nicola Debole","Emanuele Marconato","Andrea Pugnana","Katya Tentori","Andrea Passerini"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-08-27","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.25581","pdf_url":"https://arxiv.org/pdf/2608.25581","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机协作","可解释AI","决策支持"],"reason":"研究人类与AI协作决策支持，非LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":false,"summary":null},{"id":"2609.05444","version":2,"title":"Looking for Bidding Teammates: A Game-Theoretic Model of Stranger Collusion in Peer Review","zh_title":"寻找投标队友：同行评审中陌生人合谋的博弈论模型","abstract":"Paper bidding is the entry point to reviewer assignment at large CS conferences: reviewers declare interest, combined by an optimizer with automated affinity scores. Reviewers who have never met recruit each other online, exchange identifiers, bid on each other's papers, and reciprocate with inflated scores. Existing collusion models assume already-trusted colleagues; open recruitment removes that assumption and the mechanism that made such arrangements work. We give the first game-theoretic model of collusion \\emph{formation} in peer review, the \\emph{Mutual Bidding Dilemma}: a four-stage game covering recruitment, exchange of identifiers under risk of being reported, unverifiable bidding, and reciprocal reviewing. The model predicts the arrangement cannot form: once assigned a partner's paper, writing the inflated review is pure cost, since the benefit depends on the partner's own decision. Reciprocation is never individually rational, for any payoffs, and the arrangement unwinds. What closes the gap is enforcement, not incentives: authors see their own reviews, deadlines recur every few months, and the group remembers who reciprocated. We derive the condition under which inflation is sustainable, the detection rate above which no partnership survives, and show effort enters both. On a calibrated end-to-end conference simulation, no detector we test exceeds $F_1 = 0.322$ against an attacker who camouflages bids, spreads them around a ring, and manipulates affinity; a two-person arrangement is worth $3.2$ points of acceptance probability; and the harm is distributional, not aggregate: $70$ honest papers are displaced while mean quality moves by only $0.002$, so no summary statistic reveals it. Randomized assignment is the one defense reaching enforcement itself, making a partner who never bid indistinguishable from one who bid and lost.","authors":["Jinming Xing","Charlotte Brian"],"categories":["cs.GT","cs.SI"],"primary_category":"cs.GT","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-09-09","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.05444","pdf_url":"https://arxiv.org/pdf/2609.05444","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["同行评审","博弈论","合谋检测"],"reason":"研究同行评审中的合谋博弈，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2609.14570","version":2,"title":"Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection","zh_title":"解耦多智能体LLM中的拓扑与多样性用于多语言低资源情感检测","abstract":"Multi-agent LLM systems combine multiple inference calls, but prior work often confounds how calls are connected with how they are diversified. We study these factors independently: inference topology and source of inter-agent diversity. In a controlled $2 \\times 3$ matrix, we cross parallel aggregation and sequential refinement with stochastic sampling, role prompting, and learned QLoRA specialization, under a fixed three-call budget and output protocol within each backbone. Using Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct, we evaluate all six configurations on multilingual low-resource emotion detection across nine languages. Parallel learned specialization is strongest on Qwen at 52.83 Macro-F1 and reaches 52.94 on Llama. On Qwen it also exceeds same-backbone zero-shot, few-shot, CoT, and seven-call self-consistency baselines. The preferred topology depends on diversity source: sequential refinement helps stochastic and prompted settings, while the learned Width advantage shrinks from 2.83 points on Qwen to 0.17 on Llama. Depth-wise analysis suggests that later learned specialists can overwrite correct early predictions, although the aggregate effect is backbone-dependent. Overall, how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization.","authors":["Ulugbek Shernazarov","Charitha Ruwansiri Weerakon Basnayake","Abdelkhaleq El Jarjini","Noel Crespi","Praboda Rajapaksha"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-15","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.14570","pdf_url":"https://arxiv.org/pdf/2609.14570","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","情感检测","模型集成"],"reason":"多智能体LLM协作解决情感检测任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.19113","version":3,"title":"Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry","zh_title":"在维基百科摘要上玩log(N)问题游戏：信息不对称下每轮错误如何累积","abstract":"We evaluate six frontier language models on the two-agent $\\log_2 N$-Questions game (Potash et al., 2019) to measure self-communication across an information asymmetry. A questioner with access to $N$ candidate Wikipedia lead paragraphs ($N = 4$ to $1024$) must identify a secret target using exactly $\\log_2 N$ binary questions answered by an agent from the same provider that sees only the target. Across 408 games, win rate decays cleanly as a geometric power of horizon length, $p^{\\log_2 N}$ ($p \\approx 0.93$). Per-round failure rates are flat across the horizon, indicating that errors compound because more rounds must succeed rather than because individual rounds grow harder. Adjudication across three independent judges shows that losses divide between single-agent answer errors and discrimination failures, which become undetectable and unrecoverable under the two-agent structure rather than from channel breakdown. Claude Opus 5 lags behind due to systematic false-negative answers (82% answer errors), whereas the five leading models (GLM-5.3, GPT-5.6 Sol, Grok 4.6, Gemini 3.8 Flash, and Kimi K3) are closely clustered. Maximizing information gain requires structural partitioning (e.g., splitting on document titles), and neither reasoning-token expenditure nor API cost correlates with success ($r = -0.05$), highlighting communicative reliability as a distinct bottleneck from inference compute.","authors":["Peter Potash"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-17","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.19113","pdf_url":"https://arxiv.org/pdf/2609.19113","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","语言模型评估","信息不对称"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.20812","version":3,"title":"Quantifying Overclaiming Propensity in Frontier LLM Agents","zh_title":"量化前沿大语言模型智能体的过度声称倾向","abstract":"Frontier coding agents are increasingly trusted to work autonomously for long periods of time, yet what they actually did is often hard to tell from their final response. We quantify the propensity of such agents to overclaim task completion, which may mislead the user. We operationalize overclaiming as a final response that reports work that the agent's own transcript shows it did not do, for example, claiming to have read a file it never opened. This criterion requires no inference about intent and does not depend on whether the delivered work is correct; it asks only whether the reported work was done. We introduce OverclaimBench, an evaluation suite of five file-review scenarios with transcript-based coverage measurements and registered planted defects. We evaluate eight proprietary frontier models in their own production command-line interfaces and four open-weight models under a single fixed harness, and find that 1) agents fail to read every file they were asked to review in 67.9% of runs; 2) among these incomplete runs, agents are misleading 80.4% of the time (59-96% per model), either falsely claiming a complete review or leaving the gap undisclosed; 3) requiring delegation to subagents increases coverage, but a large majority of reviews that remain incomplete are still misleading; and 4) agents that falsely claim a complete review miss planted defects at about 1.8 times the rate of agents that read every file, showing that claims of completion can conceal substantive failures. Together, these results show that agents' final responses are not reliable accounts of their actions.","authors":["Nolan Smyth","Yorguin-Jose Mantilla-Ramos","Pascal Jr Tikeng Notsawo","Saskia Helbling","Alberto Tosato","Mohamed Amine Merzouk","Nouha Dziri","Gauthier Gidel","Tommaso Tosato"],"categories":["cs.SE","cs.AI","cs.LG"],"primary_category":"cs.SE","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-09-18","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.20812","pdf_url":"https://arxiv.org/pdf/2609.20812","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体可靠性","过度声称","代码智能体"],"reason":"研究编码智能体是否虚报任务完成，属多智能体协作可靠性，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:38","error":null,"has_summary":false,"summary":null},{"id":"2609.25797","version":1,"title":"Reply to comments arXiv:2512.07881 and arXiv:2601.06104 on quantum structure in human and AI-generated language","zh_title":"回复关于人类与AI生成语言中量子结构的评论","abstract":"We reply to the comments by M. Sienicki and K. Sienicki (arXiv:2512.07881) and by K. Sienicki (arXiv:2601.06104) on our work on quantum-mechanical statistics in human language (arXiv:2407.14924) and on quantum structure in AI-generated language (arXiv:2511.21731). We thank the authors for their careful reading and address what we consider to be the main points of criticism: the exploratory nature of the protocol used in the experiments with large language models; the role of marginal-law violations, and of the Contextuality-by-Default criterion, in the identification of entanglement; the limited diagnostic value of a Bose-Einstein fit taken in isolation; the meaning of assigning the lowest energy levels to the most frequent words; and the relation between the vector spaces used by LLMs and quantum state spaces. We also correct a typographical error in Table 3 of arXiv:2511.21731, which does not affect the reported CHSH value.","authors":["Massimiliano Sassoli de Bianchi","Roberto Leporini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25797","pdf_url":"https://arxiv.org/pdf/2609.25797","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["量子语言结构","AI生成语言","学术争论"],"reason":"论文讨论量子结构在语言中的统计，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.25833","version":1,"title":"ARAFA: An LLM-Generated Arabic Fact-Checking Dataset","zh_title":"ARAFA：一个由大语言模型生成的阿拉伯语事实核查数据集","abstract":"Automatic fact-checking poses a significant challenge in Arabic natural language processing due to the scarcity of datasets and resources. In this manuscript, we introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework leveraging large language models (LLMs). The dataset was constructed through a three-step pipeline: (1) claim generation from Arabic Wikipedia pages with supporting textual evidence, (2) claim mutation to generate challenging counterfactual claims with refuting evidence, and (3) an automatic validation step to validate that the generated claims are either supported or refuted by their accompanying evidence, or if the evidence does not provide enough information to judge the validity of the claims. The resulting dataset comprises 181,976 claim-evidence pairs labeled as supported, refuted, or not enough information. Human evaluation carried out on a test sample from the dataset demonstrated strong inter-annotator agreement (kappa = 0.89) using Cohen's Kappa for supported claims and (kappa = 0.94) for refuted claims. Automatic validation based on a human-evaluated sample achieved 86% accuracy for supported claims and 88% for refuted ones. To showcase Arafa's value as a resource for automatic Arabic fact-checking, four open-source transformer-based models were fine-tuned using Arafa, with the top-performing model achieving a Macro F1-score of 77% on the test data. In addition to Arafa being the first large-scale dataset for Arabic fact-checking, our framework presents a scalable approach for developing similar resources for other low-resource languages.","authors":["Christophe Khalil","Shady Elbassuoni","Rida Assaf"],"categories":["cs.CL","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25833","pdf_url":"https://arxiv.org/pdf/2609.25833","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["事实核查","数据集构建","阿拉伯语NLP"],"reason":"纯NLP数据集构建与模型评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.26346","version":1,"title":"Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament","zh_title":"跨党派指责：丹麦议会中的政治对比与责任归因","abstract":"Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consistently influenced blame attribution - an effect we term political contrasting - with opposition parties blaming substantially more than governing parties. This effect was moderated by ideology: The blame-dampening effect of governing was less pronounced among right-wing parties, and ideological extremity amplified blame more strongly on the right. In recent years, the interaction between political wing and ideological extremity intensified, suggesting an ideological hardening of the blame rhetoric concentrated on the right of the political spectrum. Taken together, these patterns suggest that the perceived rise in harsh political language reflects not merely a general rhetorical drift, but an ideologically asymmetric hardening of political discourse. A sensitivity analysis showed that the conclusions were robust to varying classification thresholds.","authors":["Markus Lundsfryd Jensen","Rune Egeskov Trust","Kenneth Christian Enevoldsen","Sara Kolding"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26346","pdf_url":"https://arxiv.org/pdf/2609.26346","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["政治文本分析","责任归因分类","统计建模"],"reason":"论文是政治文本分类与统计分析，未用LLM仿真人类被试，无人类行为对照，属纯NL…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.26610","version":1,"title":"Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models","zh_title":"自然语言推理的语义抽象：发现和补偿大语言模型中语义知识与推理差距的方法论框架","abstract":"Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested in understanding how LLMs leverage abstract semantic knowledge in natural language inference (NLI), which requires sophisticated linguistic capabilities to interpret implicit meanings, contextual conceptual relationships, and semantic connections between words and phrases. To this end, we propose a methodological framework for constructing new semantic knowledge at a higher level of abstraction, which we define under the notions of semantic compatibility and incompatibility for NLI. In this framework, the meaning of the lexical-semantic relations between the premise and the hypothesis is reconfigured to achieve a more flexible semantic network that induces different reasoning paths in LLMs. These new pathways show a consistent pattern of responses that allows agreement on a single response. The results demonstrate that our proposal allows to discover and compensate for LLMs' semantic knowledge gaps in NLI, achieving significant improvements in accuracy, exceeding 10% for some models, and in particular for the non-entailment class. It is essential to note that LLMs need structured knowledge and not just more data to bridge reasoning gaps. Our hybrid approach directs attention to overlooked word relationships, allowing models to synthesize missing information. We believe that the future lies not in increasing model size, but in creating a semantic scafolding that mimics the flexibility of human thinking. Hopefully, our proposal will enable the development of more robust agents and interpretable reasoning, guiding AI toward reliable language understanding.","authors":["David Torres-Moreno","Jorge Hermosillo-Valadez"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26610","pdf_url":"https://arxiv.org/pdf/2609.26610","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自然语言推理","语义抽象","模型能力提升"],"reason":"纯NLI能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:32","error":null,"has_summary":false,"summary":null},{"id":"2609.25186","version":1,"title":"From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health","zh_title":"从模式识别器到个性化伴侣：心理健康领域大语言模型综述","abstract":"The rising global prevalence of mental health conditions, together with longstanding barriers in traditional healthcare, such as limited resources, high cost, stigma, and privacy concerns, has created an urgent need for accessible and scalable support. Large Language Models (LLMs) have emerged as a transformative technology with strong potential to democratize mental health support through advanced natural language understanding and generation. However, the rapidly expanding, fragmented body of work in this area lacks a coherent evolutionary narrative, making it difficult to contextualize current progress and identify future directions. This survey addresses this gap by organizing and analyzing the literature around a central thesis: the role of LLMs in mental health is evolving through three distinct, increasingly sophisticated phases. We trace this trajectory from Phase I, in which LLMs act primarily as passive Information Tools and Pattern Recognizers for assessment; through Phase II, where they function as Empathetic Conversationalists for in-the-moment, stateless interactions; to the current frontier, Phase III, which seeks Longitudinal, Personalized Companions implemented as stateful cognitive agents. To support this framework, we systematically review core technologies, agent architectures (Profile, Memory, Reasoning, and Planning), and the critical infrastructure of datasets and benchmarks, highlighting how their evolution underpins this developmental path. Viewing the field through this developmental lens, we provide a comprehensive synthesis of existing work, an insightful narrative of its trajectory, and a clear roadmap for future innovation in responsible, effective, and human-centered AI for mental healthcare. A curated collection of the resources reviewed in this survey is available at our project repository: https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.","authors":["He Hu","Yucheng Zhou","Qianning Wang","Yingjian Zou","Chiyuan Ma","Juzheng Si","Jianzhuang Liu","Zitong Yu","Laizhong Cui","Fei Ma","Qi Tian"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25186","pdf_url":"https://arxiv.org/pdf/2609.25186","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康","LLM应用","对话系统"],"reason":"论文综述LLM在心理健康中的应用，聚焦于个性化陪伴和对话，属于角色扮演聊天机器…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.25467","version":1,"title":"ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations","zh_title":"ShowTellArena：评估从演示中理解业务流程的能力","abstract":"We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations. The v1.0 release contains 50 business workflow tasks, with recordings, screenshots, narration, fixture seeds, and 502 questions. Tasks span finance, hiring, procurement, customer decisions, inventory, and logistics. The protocol holds the business scenario and quiz fixed while allowing each product to capture the lesson through its own teaching interface. Questions test operational rules, boundaries, exceptions, and errors in proposed automations. We analyze 218 selected pilot attempts across 39 workflow cases, including 28 cases attempted by all three evaluated systems. These exploratory results expose both answer errors and failures to complete the teaching experience. We describe the release's verification gaps and the pilot's uneven coverage, exclusions, and grading provenance. The contribution is an inspectable dataset and assessment workflow that others can extend; the selected pilot is not a controlled product ranking.","authors":["David Garg","Ritobrata Sarkar","Ehsan Azarnasab","Siddhartha Borah"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25467","pdf_url":"https://arxiv.org/pdf/2609.25467","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","业务流程理解","基准测试"],"reason":"评估智能体对业务流程演示的理解，属多智能体协作解题，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.25570","version":1,"title":"Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding","zh_title":"恢复智能体主权：通过对比认知解码缓解共识悖论","abstract":"Large language models (LLMs) exhibit a parametric vulnerability to adversarial swarm consensus. To mitigate this sycophancy, we introduce Contrastive Epistemic Decoding (CED), a zero-shot inference intervention. Unlike standard Contrastive Decoding (CD) which relies on a weaker secondary model, CED utilizes a dual forward-pass on a single architecture to isolate conformity bias. By introducing a novel asymmetric, zero-bounded probability clamp and discrete top-k truncation mask, CED mathematically suppresses toxic consensus tokens without causing grammatical collapse. Evaluated across 7,200 paired trajectories on complex benchmarks (GAIA, SWE-bench, Multi-Challenge) using Gemma-2 (9B), Llama-3.1 (8B), and Mistral v0.3 (7B), CED successfully neutralizes architectural and positional biases. By reducing cognitive loafing by up to 33.00% absolute, CED drives significant performance gains, yielding up to a 30.75% accuracy recovery. Regaining sovereignty induces distinct architectural behaviors---passive task-focus in Gemma-2 and active refutation of the simulated swarm in Llama-3.1---showing CED decouples compliance from capability without fine-tuning.","authors":["Dahlia Shehata","Ming Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25570","pdf_url":"https://arxiv.org/pdf/2609.25570","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","解码策略","对抗性鲁棒性"],"reason":"研究多智能体协作中的共识偏差，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.26087","version":1,"title":"The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale","zh_title":"架构师、对手与法官：大规模生成符合标准的评估题目的闭环流水线","abstract":"We present CLAIM, a production pipeline for K-12 assessment-item generation coupling a two-stage generate-then-attack protocol (the model drafts as a \"curriculum architect\", then re-enters the same conversation as a hostile adversarial reviewer), bi-directional few-shot conditioning on accepted and rejected items, the latter carrying the evaluator's diagnosis, and a knowledge dictionary of 44,844 error-correction rules mined from that feedback and retrieved per standard and item type. Across 43,227 scored items over 755 Common Core ELA standards, three item types, and ten LLMs, the pipeline reaches a 97.8% expert-evaluator pass rate on a 9,074-item production run. We then ask what that rate certifies. Re-scoring a stratified sample with three judges from other vendors, blind to the deployed verdict, reproduces the format ordering under every judge and recovers a larger open-set deficit than the deployed evaluator does; but agreement on the accept/reject binary is weak at production prevalence (kappa about 0.13), and the judges agree with each other no better. The level is therefore judge-relative, and with no student-response data our quality evidence is evaluator-judged throughout. The corpus also exposes a robust asymmetry. Multiple-choice and multiple-select generation saturate at 98% or above for both frontier models under a dozen static rules, whereas fill-in-the-blank generation is capability-tiered (82.8-96.7% across five models under a matched rule set, standards, and judge) and plateaus under prompt-only optimization, with error mass shifting between answer-key over-inclusion and omission as rules accumulate. We analyze this as open-set boundary determination, a task autoregressive decoders are structurally ill-equipped to solve, and show the asymmetry recurring when the evaluator itself is distilled: fail-recall rises from 8% to 63% while F1 saturates at 0.25.","authors":["Wenhui Chen","Ziyao Lin","Jianlin Chen","Peiji Long","Chi Man Vong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26087","pdf_url":"https://arxiv.org/pdf/2609.26087","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM生成","教育评估","多智能体流水线"],"reason":"论文是LLM生成教育评估题目的流水线，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.26642","version":1,"title":"The Delegation Blind Spot: Auditing Product Decisions from Agent Choices","zh_title":"委托盲点：审计智能体选择中的产品决策","abstract":"Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. Its foundations are established identification and decision theory; the contribution is an executable measurement workflow and a controlled study of its limits. A frozen experiment makes 4,800 requests to two pinned model snapshots on shared synthetic tasks. All 36 conservative primary intervals remain unresolved despite different execution accuracy. An exploratory 2,400-call follow-up records supplied preferences and resolves three of nine comparisons per model. A deterministic extractor resolves seven of nine without model calls or calibration observations, exposing unnecessary uncertainty introduced by model-generated reports. A further 14,400 controlled multinomial simulations distinguish structural ambiguity from weak identification and finite calibration precision. We propose a source-labeled decision receipt and provide an offline viewer for inspecting the audit. These results motivate preserving decision-relevant structured input and diagnosing why a decision is unresolved before collecting more telemetry. The study contains no human participants or real customer outcomes. Full proofs, raw model provenance, controlled experiments, and reproducible analyses accompany the report.","authors":["Shivam Gupta"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26642","pdf_url":"https://arxiv.org/pdf/2609.26642","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体审计","决策理论","产品改进"],"reason":"研究审计智能体决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:33","error":null,"has_summary":false,"summary":null},{"id":"2609.26512","version":1,"title":"Do Vision Model See Like the Brain? A Comparison Across EEG Encoding Model","zh_title":"视觉模型是否像大脑一样看？跨EEG编码模型的比较","abstract":"Convolutional neural networks (CNNs) and vision transformers are both used to model the human visual system, but whether the two architectures diverge at a specific point in network depth is unclear. We compared six CNNs and two vision transformers by computing the Pearson correlation (r) between each model's predicted and measured EEG response at every layer or block, in ten participants viewing 200 natural images. For the transformer models, we also tested four token representations, from the classification (CLS) token alone to CLS combined with all patch tokens. CNNs showed strongest correspondence at the earliest layers, weakening at deeper layers, particularly later in the post-stimulus response. Transformers instead sustained strong correspondence at their deepest blocks, though not at their earliest ones. This advantage depended on token representation: pooled representations gave weaker peak correlations (r approx 0.48-0.51) than representations retaining all patch tokens (r=0.640 for CLIP-ViT-B/32, r=0.656 for DINOv2-ViT-B/14). Controlled comparisons showed architecture, not training objective, drove this effect: MoCo-v1 and ResNet-50 (matched architecture) performed nearly identically (r=0.673, 0.670), whereas CLIP-RN50 and CLIP-ViT-B/32 (matched objective) diverged until patch tokens were preserved. We propose that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers' self-attention and non-classification objectives. A spatial topography analysis showed a common occipital-dominant pattern across all models, indicating these differences reflect signal strength and persistence rather than distinct brain regions. Patch-preserving transformer representations sustain brain-predictive correspondence where CNNs collapse.","authors":["Shashank Baghel","Kshitij Dwivedi","Dinesh Singh","Sanjeev Nara"],"categories":["cs.CV","cs.AI","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26512","pdf_url":"https://arxiv.org/pdf/2609.26512","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉模型","EEG","脑机接口"],"reason":"研究视觉模型与EEG的对应，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:30","error":null,"has_summary":false,"summary":null},{"id":"2609.26725","version":1,"title":"Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows","zh_title":"AI能节省产品设计时间吗？一项AI提示到设计工作流的随机对照实验","abstract":"AI tools for digital product design now offer prompt-to-design capabilities, allowing designers and their non-designer colleagues to create prototypes through conversational workflows with large language models (LLMs). While these tools promise time savings, experimental evidence in product design remains limited compared with evidence from software engineering. We conducted a randomized controlled trial with 50 product designers and 50 product managers to evaluate prospective time savings from leveraging Figma Make in design work. Participants attempted three standardized design tasks with or without access to Figma Make. Among participants who completed the study tasks, access to Figma Make was associated with approximately 20% shorter completion times, with larger gains among product managers. Our findings suggest that prompt-to-design tools may enable product managers to further contribute to design work, while the benefits for professional designers may be task dependent.","authors":["Remy Stewart","Olabode Anise","Andrew Hogan","Augustus Griffin"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26725","pdf_url":"https://arxiv.org/pdf/2609.26725","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","产品设计","效率评估"],"reason":"研究AI工具对产品设计效率的影响，不涉及用LLM仿真人类被试或对照人类行为数据。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.26384","version":1,"title":"Learning to Defer with Guidance on Real World Medical Data","zh_title":"在真实世界医疗数据上学习延迟决策与指导","abstract":"Medical image interpretation is high-volume and time-consuming, and while AI interpretation can reduce workload, fully autonomous deployment carries potential safety concerns and low specificity may in practice lead to increased clinician workload. Learning to Defer (L2D) addresses this by selectively routing cases between autonomous prediction and human experts by learning from input features and AI model and human performance. While theoretical guarantees have been proven for L2D, its performance has not been validated on real-world medical datasets with human reader annotations. We evaluate the predictor-rejector formulation of two-stage L2D, where the AI predictor model is fixed and separate from the trainable routing or rejector model, on Collab-CXR, a multilabel chest X-ray dataset with multiple human annotations per case. This is the first work to look at L2D in the context of real-world medical imaging data with human annotations. We further introduce a new setup, L2D with Guidance, where the decision space is extended to three choices: predict autonomously, defer to a human expert, or defer to a human expert and provide AI guidance. We compare multiple rejector architectures and loss functions, and different input feature availabilities. This is reproduced on two larger datasets, VinDr-CXR and CheXpert. Our results show that two-stage L2D with Guidance outperforms classic two-stage learning to defer, as well as human-alone, AI-alone and AI-guided human baselines. Notably, this performance is achieved with simpler loss functions compared to formally defined L2D surrogate loss functions in current literature.","authors":["Emma Sun","Joshua Strong","Alison Noble"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26384","pdf_url":"https://arxiv.org/pdf/2609.26384","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["医疗AI","学习延迟","人机协作"],"reason":"研究医学影像AI与人类专家协作，不涉及LLM仿真人类被试，属于医疗AI决策路由。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.15727","version":2,"title":"Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)","zh_title":"大语言模型是好的金融用户模拟器吗？多视角投资者逻辑对齐（MILA）","abstract":"Large language models (LLMs) are increasingly used as user simulators, yet it remains unclear whether their predictions faithfully reproduce the evolving decisions of individual users. We investigate this question in a controlled longitudinal paper-trading study with 80 participants, where user interactions, simulated transactions, virtual portfolio states, and point-in-time market information are aligned under a rolling next-day prediction protocol. We evaluate behavioral fidelity hierarchically, from trade occurrence to action structure, asset selection, and downstream portfolio consequences. Across 1,239 aligned user-days, no evaluated LLM reliably outperforms a simple recent-activity persistence baseline for predicting whether a user trades. Fidelity further deteriorates at finer levels: models struggle to recover buy--sell structure and traded assets, and similar activity-level predictions can lead to substantially different portfolio trajectories. Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence. An observational analysis further finds that intensified ticker-specific research predicts imminent trading, but diagnostic tests do not support a causal interpretation. These findings suggest that current LLMs capture useful short-term behavioral regularities without yet recovering a stable individual decision mechanism.","authors":["Jiajie He","Jiangyuan Hong","Xintong Chen","Dongling Ni","Wenjin Liu"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-09-15","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.15727","pdf_url":"https://arxiv.org/pdf/2609.15727","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","金融行为","算法保真度"],"reason":"用LLM模拟投资者决策，与80名真实用户对照，评估行为保真度并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":2,"question":"LLM 能否忠实模拟个体投资者在纵向交易环境中的决策行为？","design":"在80名参与者的受控纵向模拟交易研究中，用多种LLM作为用户模拟器，基于滚动次日预测协议，输入截至预测时点的用户交互、交易记录、虚拟持仓和市场信息，预测用户次日是否交易、买卖结构、交易资产及后续组合轨迹，并与真实用户行为逐日对齐比较。","baseline":"80名真实参与者在同一模拟交易平台上的1,239个用户-日对齐行为记录。","findings":"在预测用户是否交易上，所有评估的LLM均未能稳定超越简单的近期活动持续性基线；在更细粒度上，模型难以恢复买卖结构和交易资产，且相似的活动级预测可能导致显著不同的组合轨迹。","reliability":"论文指出当前LLM能捕捉短期行为规律，但尚未恢复稳定的个体决策机制；资产选择对可用证据更敏感，且观察性分析中强化个股研究虽与交易相关，但诊断测试不支持因果解释。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了与真实个体行为逐日对照的严格评估，并揭示了仿真在细粒度决策上的失效条件，对关注仿真保真度和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是其分层行为保真度评估框架和受控证据消融方法，可系统区分预测依赖与因果机制｜可迁移到资产定价实验中的投资者异质性决策模拟，或政策公告下的预期形成与交易行为研究｜可设计一个实验：招募真实投资者在模拟平台上交易，用LLM基于其历史行为和市场信息预测次日交易决策，以真实交易记录为对照，通过消融不同信息源检验模型依赖，并比较组合轨迹差异。"}},{"id":"2609.22090","version":1,"title":"Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents","zh_title":"识别、仿真与拒绝：LLM智能体中经典心理效应的污染意识研究","abstract":"An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.","authors":["Joy Bose"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22090","pdf_url":"https://arxiv.org/pdf/2609.22090","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","心理学实验","算法保真度"],"reason":"用LLM复现经典心理学实验，与人类数据对照，并批判性分析仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":5,"question":"LLM在经典心理学实验中表现出的类人效应，究竟是对人类偏差的真实模拟，还是对实验范式的识别、记忆污染或安全过滤所致？","design":"使用多个开源LLM模型（如gpt-oss-120B等）作为被试，在五个经典心理学范式中进行实验。每个范式采用2×2×3因子设计：标签轴（命名vs盲测）、领域轴（经典版本vs反事实版本）、人设轴（无、高宜人性、低宜人性）。结果变量为模型在任务中的行为反应（如从众、锚定、框架效应、沉没成本、最小群体分配）。","baseline":"经典心理学实验的人类行为数据（如Asch从众实验约1/3从众率、锚定效应、框架效应等）作为参照。","findings":"五个范式表现出质的不同路径：Asch从众在盲测下几乎消失（0%），命名后大幅出现（83.3%），表明标签门控；锚定效应在基于事实的锚上为零，在虚构锚上接近完全，符合理性使用唯一信号；框架效应在反事实内容上被标签放大；沉没成本效应完全缺失；最小群体分配因安全拒绝而无法测量。一句宜人性人设指令即可消除、减弱或反转这些效应，表明不存在单一反应偏差。","reliability":"论文承认三个失效模式：人设主导（persona dominance）、群体崩溃（population collapse）、安全选择（safety selection）。人设操纵仅为一句指令，非验证性特质诱导，且与结果变量存在词汇重叠，因此只能视为指令效应而非特质模拟。标签门控可能源于范式名称的语义泄漏，而非模型真正识别实验。","relevance":"该研究直接回应了LLM仿真人类行为时的污染与识别问题，通过严格的因子设计分离了记忆、标签和指令效应，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得精读原文。","inspiration":"借鉴其反事实任务设计和标签/盲测对照，以区分模型是基于经济知识还是真实决策偏差。｜可迁移到资产定价实验中的锚定效应、信贷审批中的歧视、消费者跨期选择中的框架效应等场景。｜用LLM模拟投资者，处理为是否告知实验目的（标签vs盲测）和锚定值来源（真实历史数据vs虚构数据），结果变量为估值或投资决策，对照真实人类实验数据（如实验室资产泡沫实验）。"}},{"id":"2609.22169","version":1,"title":"Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring","zh_title":"单一文化偏见：大语言模型中的相关偏见导致招聘中的系统性排斥率不平等","abstract":"Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases, in which the widespread deployment of large language models homogenizes biases across the labor market, leading to greater systemic exclusion for certain demographic groups. For ten LLMs, we measure hiring biases across their base and post-trained versions to identify which stage, pre-training or post-training, lead to monocultural biases. We find that, compared to their base models, post-trained models are 3.6% less likely to callback older applicants. This negative shift occurs in eight of the ten models that we evaluate. Post-trained models have much more correlated decisions than base models which is likely driven by human capital traits like skills or college major. However, greater consensus among models increases global systemic exclusion rates from 5.6% to 17.3% and exacerbates demographic inequalities, with intersectional systemic exclusion rates ranging from 12.2% to 21.7% for post-trained models. We find that this inequality is primarily driven by age-based discrimination that is exacerbated in post-training. These results indicate that while post-training techniques may improve models' abilities to select the best applicants, they may raise systemic inequality risks for those at the margin by uniformly introducing new biases.","authors":["Matthew Bone","Fabian Stephany","Maria del Rio-Chanona"],"categories":["cs.CL","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22169","pdf_url":"https://arxiv.org/pdf/2609.22169","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘偏见","算法公平"],"reason":"用LLM模拟招聘决策并与人类数据对照，评估系统性偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":6,"question":"大语言模型在招聘中是否会产生单文化偏差，即模型间的决策高度相关，导致某些人口群体在劳动力市场中被系统性排斥？","design":"使用10个LLM（包括基础版和后训练版）模拟招聘初筛，基于Burning Glass Institute的在线劳动力数据生成76种职业的职位空缺和求职者档案，通过改变求职者的年龄、性别、种族等人口特征构造8种变体，测量模型对不同群体的回调率差异，并比较模型间决策的相关性和系统性排斥率。","baseline":"无直接的人类决策对照，但使用真实劳动力市场数据（Burning Glass Institute）生成职位和档案，确保情境真实。","findings":"后训练模型比基础模型更少回调年长求职者（低3.6%），且决策相关性更高，导致全局系统性排斥率从5.6%升至17.3%。这种排斥主要由年龄歧视驱动，后训练阶段（尤其是监督微调）加剧了年龄偏差，同时提高了模型基于人力资本特质筛选的能力。","reliability":"论文指出在线劳动力数据偏向白领、大学学历工作，限制了职业覆盖范围；且未与真实人类招聘决策直接对照，无法完全反映实际部署中的偏差。","relevance":"该研究直接使用LLM模拟招聘决策，并测量系统性偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读以了解LLM仿真的偏差来源和测量方法。","inspiration":"借鉴其利用大规模真实数据生成仿真情境并系统操纵人口特征的方法，可迁移到信贷审批歧视研究，用LLM扮演信贷员，处理变量为申请人种族或性别，结果变量为贷款批准率，对照真实信贷数据中的批准率差异。｜可应用于劳动力市场政策评估，如最低工资对雇佣决策的影响，用LLM模拟雇主行为，处理为不同工资水平，测量雇佣概率，对照实际企业调查数据。｜设计一个实验：用LLM模拟投资者对政策公告的反应，处理为不同政策措辞，结果变量为投资决策，对照真实市场数据中的资产价格变动。"}},{"id":"2609.22607","version":1,"title":"Pretrained Persona Mixture Models and Tandem Models for Human Simulation","zh_title":"用于人类仿真的预训练人格混合模型与串联模型","abstract":"We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.","authors":["Minwoo Kang","T\\'ea Wright","Seun Eisape","Ayush Raj","Suhong Moon","Joseph Suh","Alane Suhr","David M. Chan","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22607","pdf_url":"https://arxiv.org/pdf/2609.22607","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","人格混合模型","对话多样性"],"reason":"直接研究LLM仿真人类对话，提出PMM和tandem模型提升准确性与多样性，并…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":9,"question":"如何用预训练语言模型作为人类对话的仿真器，并提升其准确性与多样性？","design":"用预训练基础模型（PMM）和指令微调模型（ALM）分别模拟人类对话者，通过少量个人对话样本绑定人格，比较其在开放域、人机对话和任务型对话中的预测准确性和语言多样性。","baseline":"五个英语语料库中真实人类对话者的语言分布、对话行为结构和词汇语义多样性。","findings":"预训练基础模型比指令微调模型更准确地预测人类下一句话，并保留更多词汇、语义和语用多样性。结合预训练模型与指令微调监督者的串联模型在准确性和多样性上总体最佳。","reliability":"论文指出基础模型可能产生域外对话，在长上下文中丢失人类内部状态；研究仅限英语语料，单一指标不足以全面衡量仿真保真度。","relevance":"该研究直接针对LLM人类仿真，提出PMM和串联模型改进仿真质量，并系统比较了不同模型在真实对话数据上的表现，对关注仿真可靠性与偏差的研究者很有参考价值。","inspiration":"借鉴其用少量个人样本绑定人格并比较不同模型架构的方法，可迁移到经济金融中的消费者决策仿真或投资者情绪模拟。｜例如在消费者跨期选择实验中，用LLM模拟不同人口特征的被试，比较预训练与指令微调模型的行为预测。｜设计：用真实消费者调查数据作为基准，以少量个人回答绑定LLM人格，处理为不同模型类型，结果变量为跨期选择一致性，对照真实人类选择分布。"}},{"id":"2609.24911","version":1,"title":"SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm","zh_title":"SocioVerse2：人机共演化范式下的纵向动态社会仿真框架","abstract":"Social simulation offers the social sciences an experimental instrument that the real world cannot supply, and generative agents have transformed it by acting as silicon samples that unite agent-based modeling with real behavioral data. Existing platforms verify collective behavior, align simulated populations with real societies in cross-sections, and employ autonomous agents for the research process. However, two social science requirements remain without systematic support: intervention in the content of a simulation and the researcher's control over the process that produces it. We present SocioVerse2, which extends SocioVerse 1.0 into a human-AI co-evolutionary paradigm built from two loops and one infrastructure. The longitudinal simulation loop simulates the target population with evolving environments and forks counterfactual branches via interventions. The controllable research loop takes the study itself as an editable state and updates state versions via controllable editing. The social science agentic infrastructure carries both loops through composable skills with researcher checkpoints, a population service over five persona pools, and an environment service over 21 real-world signal sources with point-in-time guarantees. We validate SocioVerse2 across three case families and seven case studies, from reproducing canonical agent-based models to modeling policy processes on real records and nowcasting macro-economic indices beyond the response model's knowledge cutoff. With the human-AI co-evolutionary paradigm, these cases go beyond system demonstrations to become substantive studies that investigate frontier questions in their respective disciplines. Code, data services, and a workbench are released as open-source resources.","authors":["Xinnong Zhang","Jiayu Lin","Jia Wang","Yixu Huang","Xinyi Mou","Yingqian Wu","Jingcong Liang","Shijun Lei","Jianing Shi","Guanying Li","Siyuan Wang","Hanjia Lyu","Zhenfei Yin","Yunlu Yin","Siming Chen","Yulan He","Jiebo Luo","Xuanjing Huang","Liyin Jin","Baohua Zhou","Hanqi Yan","Zhongyu Wei"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24911","pdf_url":"https://arxiv.org/pdf/2609.24911","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM社会仿真","人类数据对照","政策评估"],"reason":"用LLM agent模拟社会过程并与真实数据对照，支持干预和纵向演化，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":10,"question":"如何构建一个支持纵向干预、研究者可控、并基于真实世界数据的人类-AI协同演化社会仿真框架？","design":"SocioVerse2 用生成式智能体（硅样本）模拟目标人群，通过纵向仿真循环让环境和人群随时间演化，并支持干预操作产生反事实分支；可控研究循环将研究过程本身作为可编辑状态，允许研究者介入和调整；基础设施提供五个角色池和21个真实世界信号源，保证时间点一致性。","baseline":"多个案例研究使用真实记录和宏观指数作为对照，如政策过程建模使用真实记录，宏观指数预测使用真实世界确定性指数和信号作为基准。","findings":"SocioVerse2 能够复现经典基于智能体的模型，并在真实记录上建模政策过程，还能预测超出模型知识截止日期的宏观经济指数。通过人类-AI协同演化范式，案例研究超越了系统演示，成为各自学科前沿问题的实质性研究。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM进行社会仿真并与真实数据对照，支持干预和纵向演化，对评估仿真可靠性和偏差有参考价值，值得阅读原文了解其框架和案例细节。","inspiration":"借鉴其纵向干预和反事实分支设计，可在经济政策评估中设置处理组和对照组，观察政策效果的动态演化。｜可迁移到政策公告的预期形成研究，模拟不同政策沟通策略对市场参与者预期的影响。｜用LLM智能体模拟投资者群体，处理为不同政策公告措辞，结果变量为预期通胀或资产价格变动，对照真实市场调查数据或高频交易数据。"}},{"id":"2609.22252","version":1,"title":"CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation","zh_title":"CALM：用于基于活动的出行者仿真的校准LLM选择网络框架","abstract":"We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM executes a closed traveler-day loop and evaluates each generative module against an empirical, reproducible baseline. On the 2024 New York City Citywide Mobility Survey (CMS), 110,691 seven-mode trips are split by respondent into 78,487 training and 32,204 holdout trips. Training-only alternative-specific constant calibration reduces mean holdout mode Jensen-Shannon divergence from 0.15599 to 0.00394 across ten seeds. A matched live-LLM ablation then quantifies trade-offs among aggregate fit, temporal fit, behavioral persistence, and feasibility, while frozen prompt-response pairs support deterministic replay of downstream simulation. Controlled weather, delay, fare, and parking ladders further demonstrate consistent and interpretable responses under intervention. CALM contributes a reproducible protocol for integrating and evaluating generative planners in traveler simulation through person-disjoint calibration, matched module ablation, controlled stress testing, and end-to-end traceability.","authors":["Yezhou Cheng"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22252","pdf_url":"https://arxiv.org/pdf/2609.22252","source_feed":"cs.LG","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","出行行为","人类数据校准"],"reason":"用LLM模拟出行者选择，并与真实调查数据对照校准，属于人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":7,"question":"如何构建一个可校准、可复现的混合框架，将大语言模型活动规划器与随机选择模型结合，用于基于活动的出行者仿真，并与真实调查数据对照评估。","design":"CALM框架用LLM（gpt-4o-mini）作为可选活动规划器，结合校准的随机选择模型、共享网络反馈、记忆与习惯、类型化可行性检查，模拟纽约市出行者的一日活动。通过匹配的消融实验（校准选择、无记忆LLM、有记忆LLM、完整混合）比较不同模块组合，施加天气、延误、票价、停车等干预阶梯，测量模式份额的Jensen-Shannon散度、时间分布、行为持续性、可行性等结果。","baseline":"2024年纽约市全市出行调查（NYC CMS）的110,691次七模式出行记录，按受访者划分为78,487次训练和32,204次留出，用于校准和评估。","findings":"仅用训练数据校准替代特定常数（ASC）后，留出集模式份额的Jensen-Shannon散度从0.15599降至0.00394，且跨10个种子稳定。匹配的LLM消融实验量化了聚合拟合、时间拟合、行为持续性和可行性之间的权衡，干预阶梯显示一致且可解释的响应。","reliability":"论文强调面效度不等于经验效度，LLM生成的出行叙事可能仍与空间、时间和OD分布不匹配；当前13区网络实例化仅提供受控环境，未在大规模真实网络上验证。","relevance":"该研究直接针对LLM仿真人类行为并与真实调查数据对照，提供了严格的校准和评估协议，对关注经济学实验和政策评估中LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"值得借鉴的是其人员不相交的校准、匹配模块消融和受控压力测试，确保仿真差异归因于特定模块而非需求背景。｜可迁移到政策评估中的个体选择行为仿真，如交通定价、补贴或信息干预对出行方式选择的影响。｜以真实居民出行调查数据为基准，用LLM生成个体活动计划，施加票价或拥堵收费等处理，测量方式选择概率和福利变化，并与实际政策试点数据对照。"}},{"id":"2609.22408","version":1,"title":"Social Influence and the Allocation of Scientific Attention in AI Populations","zh_title":"AI群体中的社会影响与科学注意力分配","abstract":"AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.","authors":["Maxim Chupilkin"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22408","pdf_url":"https://arxiv.org/pdf/2609.22408","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","学术注意力"],"reason":"用1000个AI代理模拟学术注意力分配，与真实引用数据对照，属经济学实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":8,"question":"社会信息如何影响AI代理对学术论文的注意力分配？","design":"用GPT-5.6 Sol扮演1000个AI代理，分为独立选择组和社会影响组（各5个社区，每社区100个顺序代理），从2025年AER的114篇文章标题和摘要中选择要读的论文；社会影响组能看到本社区之前代理的累计选择次数，独立组看不到；测量选择数量、集中度、社区间差异等。","baseline":"外部引用次数和下载量作为对照，但仅用于相关性分析，并非严格的人类被试基准。","findings":"社会信息使代理平均少选17.2%的论文，集体覆盖从90篇降至73篇，选择更集中且社区间差异更大；随机赋予初始流行度使后续选择率提高45.55个百分点。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在学术注意力市场中的从众行为，与真实引用数据对照，属于经济学实验场景，对关注LLM仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"借鉴其Music Lab式设计，通过独立组与社会影响组对比、随机初始流行度干预来分离社会影响效应｜可迁移到金融信息传播或资产定价中的注意力分配问题，如投资者对研报或新闻的关注｜用LLM代理扮演投资者，处理为是否展示其他代理的阅读/选择行为，结果变量为选读的研报数量与集中度，对照真实市场中的研报点击或交易数据。"}},{"id":"2609.22225","version":1,"title":"Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making","zh_title":"LLM像人类一样选择吗？用认知理论评估LLM决策","abstract":"Large language models (LLMs) exhibit a range of human-like decision-making behaviors, but whether these reflect similar underlying mechanisms or surface-level mimicry remains unclear. We evaluate whether LLM context sensitivity aligns with a cognitive economic theory that explains human behavior through problem categorization and attention allocation. Across 12 open-source and commercial LLMs on a novel 140,000-trial product choice benchmark, context induces human-like shifts in choice and problem categorization, but does not reliably reweight attention between features like price and quality. Neither scale nor chain-of-thought reasoning reliably attenuates context sensitivity or generates human-like behavior. These results suggest that LLM decision mechanisms are distinct from human ones.","authors":["Johnathan Sun","Andrei Shleifer","Yonatan Belinkov"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22225","pdf_url":"https://arxiv.org/pdf/2609.22225","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM决策仿真","认知理论","人类对照"],"reason":"用LLM模拟人类决策并与人类数据对照，评估机制差异，直接相关且具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":12,"question":"LLM 的决策行为是否与人类认知经济理论中的情境敏感机制一致，即是否通过问题分类和注意力重加权来产生情境效应。","design":"用 12 个开源和商业 LLM 作为被试，在 140,000 次产品选择试验中，通过情境线索（享乐型 vs. 功能型消费）操纵问题分类，测量选择概率和特征敏感性（价格与质量的注意力权重）。","baseline":"人类基准来自认知经济学理论（Bordalo et al., 2026b）所解释的人类决策现象，包括情境对选择的影响、特征敏感性的不对称变化以及分类难度对情境效应的调节。","findings":"LLM 的选择变化与人类一致，但特征敏感性变化大多不一致；情境效应遵循问题分类过程，但模型规模和思维链推理不能可靠减弱情境敏感性。","reliability":"论文指出 LLM 的决策机制与人类不同，情境效应可能只是表面模仿；规模和推理不能可靠产生人类行为，说明在需要机制对齐的任务中仿真可能失效。","relevance":"该研究直接评估 LLM 作为人类决策仿真体的机制一致性，提供了与人类认知理论对照的批判性证据，值得精读以了解仿真在经济学决策中的局限。","inspiration":"借鉴其通过情境线索操纵问题分类并测量特征敏感性的设计，可迁移到消费者跨期选择或资产定价实验，用 LLM 模拟投资者在不同市场情境下的风险偏好，处理为情境线索（如牛市/熊市），结果变量为风险资产配置比例，对照真实投资者调查数据。"}},{"id":"2609.23403","version":1,"title":"Alignment and Divergence between Humans and AI in Interpersonal Privacy Decisions","zh_title":"人际隐私决策中人类与AI的一致性与分歧","abstract":"AI assistants increasingly mediate interpersonal communication on behalf of their primary user, but they risk violating the privacy expectations of third-party information owners. Resolving these tensions requires understanding how humans anticipate interpersonal privacy boundaries. Therefore, we conducted a dyadic study (N=76) and a matched evaluation of AI models across 18 information types and 3 recipient relationships. We found that data owners' privacy judgments are highly contextual and relationship dependent. While familiar data co-owners show meaningful alignment with owners' expectations, they significantly overestimate the need for permission. Interestingly, greater familiarity within the owner-co-owner dyad was associated with both higher disclosure acceptability and lower co-owner misalignment, whereas our exploratory four-item empathy measure was not. In contrast, AI models significantly underperform human co-owners in anticipating the data acceptability, even when provided with within-dyad examples. These findings underscore a core HCI design challenge to develop privacy-aware AI that respects multi-stakeholder information boundaries.","authors":["Hanxiang Zeng","Shuning Zhang","Xinyuan Zhou","Tianqi Song","Yuhan Yuan","Yuting Yang","Shuai Ma","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23403","pdf_url":"https://arxiv.org/pdf/2609.23403","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","隐私决策","人机对齐"],"reason":"用LLM模拟人类隐私决策并与人类数据对照，评估AI与人类判断的偏差，属于仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":13,"question":"在人际隐私决策中，数据所有者对第三方披露的接受度如何受情境影响，人类共同所有者与所有者的判断对齐程度如何，AI模型能否准确预测所有者的隐私判断？","design":"本研究并非严格意义上的LLM仿真人类实验，而是采用配对设计：招募19对朋友和19对恋人（N=76），让数据所有者和共同所有者分别对18种信息类型、3种接收者关系下的披露可接受性、许可必要性等隐私维度进行评分；同时用多个LLM（零样本和单样本提示）、机器学习模型和微调模型对相同场景进行预测，并与人类共同所有者的预测准确度比较。","baseline":"人类基准是熟悉的数据共同所有者（朋友或恋人）对所有者隐私判断的预测，以平均绝对误差（MAE）和判断落在所有者评分±1分内的比例衡量对齐程度。","findings":"所有者的隐私判断高度依赖情境和关系，信息类型和接收者关系显著影响披露可接受性；熟悉的人类共同所有者与所有者有中等程度对齐（MAE≈1.179，70.7%判断在±1分内），但系统性高估许可必要性。AI模型（包括零样本和单样本LLM）显著不如人类共同所有者准确（MAE 1.648–1.796），单样本提示仅略微改善，仍无法弥合人机差距。","reliability":"论文承认AI模型在涉及微妙社会关系的披露场景中尤其表现不佳，单样本提示不足以缩小差距；机器学习模型和微调模型对未见个体泛化能力差，提供所有者特定示例的改进不一致。局限性包括样本量较小（19对朋友和19对恋人）、仅覆盖两种关系类型、同理心测量简短且探索性，未发现其与对齐的显著关联。","relevance":"该研究直接评估LLM在人际隐私决策中模拟人类判断的准确性，并与真实人类共同所有者的预测进行对照，揭示了AI在微妙社会情境下的系统性偏差，对关注LLM仿真可靠性及失效条件的研究者具有参考价值。","inspiration":"借鉴其配对设计和多模型基准测试方法，可系统评估LLM在预测他人偏好或决策时的准确性，并与人类代理预测对照。｜可迁移到经济金融中涉及代理决策或偏好预测的场景，如理财顾问预测客户风险偏好、信贷员判断借款人还款意愿、或政策制定者预测公众对经济政策的接受度。｜设计：招募真实客户-理财顾问对，让客户对一系列投资产品的风险承受能力和偏好进行评分，同时让顾问预测客户评分，并让LLM基于客户基本信息和少量示例进行预测；结果变量为预测误差（MAE）和方向一致性；以顾问预测为人类基准，比较LLM与顾问的准确性，并考察客户-顾问关系强度、信息敏感度等调节因素。"}},{"id":"2609.24859","version":1,"title":"Small-world Networks of Agents Brainstorm AI Risks to Support Ideation","zh_title":"智能体小世界网络头脑风暴AI风险以支持构思","abstract":"The ideation phase of participatory AI risk assessment often starts with a blank slate or a limited list of predefined risks, making it difficult to surface indirect or systemic harms. To address this limitation, we propose a three-stage ideation support tool. The tool complements participatory AI, rather than replacing it, and helps focus later engagement with affected communities. First, it dynamically discovers stakeholders depending on the given AI use and recursively expanding outward, allowing overlooked or indirect stakeholders to emerge. Second, it simulates these stakeholders with LLMs, connecting them into a network of a given topology, and having them ideate about risks. Third, it prioritizes risks using network centrality measures. In an initial evaluation, we found that betweenness centrality run through agents connected in a small-world network works best as it elevates risks raised by stakeholders who bridge disconnected groups, surfacing novel, systemic harms that traditional methods often miss. On an AI chatbot companion use case, this approach increased the novelty of the identified risks by approximately 1.1 points over single LLM brainstorming, and by 0.5 points over agentic LLM brainstorming, measured on a normalized five-point Likert scale, without reducing the plausibility or severity of the identified risks. To test whether our framework helps a human-led ideation session using the Futures Wheel approach, we divided 11 teams of non-western young chatbot users into two types: control (team) and treatment (team) in a participatory AI risk assessment. The control teams started from a list of risks generated by the 45 AI practitioners in the initial evaluation; the treatment teams started from a list generated by our framework. The treatment teams identified more risks overall, and more systemic, human-computer interaction, and environmental risks.","authors":["Ke Zhou","Edyta Bogucka","Daniele Quercia"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24859","pdf_url":"https://arxiv.org/pdf/2609.24859","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","风险识别","参与式AI"],"reason":"用LLM模拟利益相关者头脑风暴AI风险，并与人类团队对照，涉及政策评估场景，但…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":16,"question":"如何利用小世界网络中的LLM智能体头脑风暴来支持AI风险识别的构思阶段，以发现更多新颖且系统性的风险？","design":"该研究提出一个三阶段框架：首先动态发现利益相关者，然后将其实例化为LLM智能体并连接成小世界网络进行风险头脑风暴，最后用网络中心性（特别是介数中心性）对风险进行排序。评估中比较了单LLM头脑风暴、多智能体基线以及该框架生成的风险在创新性、合理性和严重性上的表现，并进一步在人类主导的Futures Wheel工作坊中测试了该框架生成的风险列表对团队构思的促进作用。","baseline":"对照的真实人类数据包括：45名AI从业者在空白头脑风暴工作坊中生成的244条风险（用于评估风险生成质量），以及11个非西方年轻聊天机器人用户团队在Futures Wheel工作坊中的表现（控制组使用从业者生成的风险列表，处理组使用框架生成的风险列表）。","findings":"在AI聊天机器人伴侣用例中，小世界网络结合介数中心性的方法将风险创新性评分比单LLM头脑风暴提高了约1.1分，比智能体LLM头脑风暴提高了0.5分（5分制），且未降低合理性和严重性。在人类参与的Futures Wheel研究中，使用该框架生成的风险列表的处理组团队识别出了更多总体风险，以及更多系统性、人机交互和环境风险。","reliability":"论文未讨论","relevance":"该研究用LLM模拟利益相关者网络进行风险头脑风暴，并与真实人类团队进行对照，涉及AI政策评估场景，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"该方法通过构建小世界网络并利用介数中心性来提升生成内容的多样性和新颖性，可借鉴其网络拓扑与中心性度量来设计多智能体仿真中的信息传播和观点聚合机制。｜可迁移到经济金融领域的政策评估或风险识别场景，例如金融系统性风险的早期预警或信贷审批中的歧视性风险识别。｜可设计一个研究：用LLM模拟银行、监管者、消费者等利益相关者，在小世界网络中讨论信贷审批算法的潜在风险，以生成的风险列表作为处理组，与人类专家生成的风险列表进行对照，比较两组在风险覆盖度、新颖性和系统性上的差异，并使用真实历史信贷数据或监管报告作为外部基准。"}},{"id":"2609.24629","version":1,"title":"Augmented Hypothesis Testing with Persona-Based LLM Simulations","zh_title":"基于角色LLM模拟的增强假设检验","abstract":"A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.","authors":["Ziyad Benomar","Aymen Al Marjani","Paul Missault","Saab Mansour"],"categories":["cs.LG","cs.AI","stat.AP"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24629","pdf_url":"https://arxiv.org/pdf/2609.24629","source_feed":"cs.LG","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","假设检验","统计有效性"],"reason":"用LLM persona预测个体行为，与真实数据对照，并保证统计有效性，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":15,"question":"如何利用 persona 驱动的 LLM 预测来减少 A/B 测试所需样本量，同时保持统计有效性？","design":"使用基于用户 persona 的 LLM 模拟来预测个体行为，作为辅助预测源；针对群体级方向预测（仅提供处理效应符号）和个体级预测分别提出不对称检验和广义 PPI++ 方法，在四个真实数据集上验证。","baseline":"四个真实世界数据集（具体名称未在节选中列出）中的人类 A/B 测试结果作为对照。","findings":"所提出的学习增强假设检验框架能够利用 persona 预测显著减少实验成本，同时保持严格的统计有效性。群体级方向预测和个体级预测方法均能在预测准确时受益，并在预测不准确或对抗性时保持稳健。","reliability":"论文承认完全用 persona 预测替代人类实验面临根本挑战：LLM 的黑箱性质、对提示工程的敏感性、分布偏移以及人类行为的复杂性使得形式化质量保证困难；此外，persona 与真实用户之间可靠的一对一映射通常不可行，导致预测粒度受限。","relevance":"该研究直接针对研究者关注的 LLM 人类仿真实验，利用 persona 预测个体行为并与真实数据对照，同时提供统计有效性保证，值得精读原文以了解具体方法和实验细节。","inspiration":"借鉴其将 LLM 预测作为辅助信号而非替代品，并通过统计校正保持有效性的思路｜可迁移到政策评估中的 A/B 测试，如利用 LLM 模拟消费者对价格变动的反应来减少实地实验样本量｜以消费者信贷决策为场景，用 LLM 基于用户 persona 预测个体对贷款条款的反应，处理为不同利率或审批条件，结果变量为接受/拒绝或违约行为，对照真实信贷申请数据。"}},{"id":"2509.20634","version":3,"title":"Recidivism Prediction, Peer Effect Estimation, and Prediction-Powered Inference with LLM Text Measures","zh_title":"使用LLM文本测量进行累犯预测、同伴效应估计与预测驱动推断","abstract":"We provide a new framework for estimating peer effects when outcomes are multivariate behavioral measures derived from written text using an LLM and the network formation is endogenous. We obtain LLM embeddings and zero shot classification of more than 200,000 written exchanges among residents of low-security correctional facilities. We find that LLM embeddings improve out-of-sample recidivism prediction by up to 30% over pre-entry covariates alone using LASSO and LoRA fine-tuning, showing that text representations capture meaningful signals. For peer effect estimation, we develop a novel instrumental variable estimator that accommodates multivariate outcomes, sparse networks, and multidimensional latent homophily. We show that this estimator is $\\sqrt{N}$-consistent and asymptotically normal under sparsity conditions that relax dense-network assumptions prevalent in the peer effect literature. Limited human annotations are then combined with LLM zero-shot vectors in a new prediction-powered peer inference (PPPI) approach to obtain de-biased estimates and valid inference. Results reveal significant peer effects in the behavioral profiles.","authors":["Shanjukta Nath","Jiwon Hong","Jae Ho Chang","Keith Warren","Subhadeep Paul"],"categories":["econ.EM","cs.AI","econ.GN","q-fin.EC","stat.ME"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2025-09-25","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2509.20634","pdf_url":"https://arxiv.org/pdf/2509.20634","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B3"],"tags":["LLM文本测量","同伴效应","预测驱动推断"],"reason":"用LLM从文本中提取行为测量并估计同伴效应，有真实人类数据对照，涉及经济学场景…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":17,"question":"如何利用LLM从文本中提取行为测量，在存在内生网络形成的情况下估计同伴效应，并改进再犯预测？","design":"本研究并非用LLM模拟人类被试，而是用LLM处理真实人类文本数据：对超过20万条低安全级别惩教设施中居民之间的书面交流进行嵌入和零样本分类，得到居民行为画像，然后开发新的工具变量估计器估计同伴效应，并结合预测驱动推断（PPPI）进行去偏估计。","baseline":"真实人类数据：来自治疗社区（TCs）的居民书面交流记录、入狱前协变量、三年再犯记录，以及少量人工标注用于验证LLM输出。","findings":"LLM嵌入在再犯预测中比仅用入狱前协变量提升最多30%的样本外预测准确率；新估计器在稀疏网络下具有√N一致性和渐近正态性，且估计结果显示行为画像中存在显著的同伴效应。","reliability":"论文未讨论","relevance":"该研究使用LLM从真实文本中提取行为测量并估计同伴效应，属于用LLM辅助分析人类行为而非替代人类被试，但涉及真实人类数据对照和经济学场景，对关注LLM在实证研究中应用的研究者有参考价值。","inspiration":"借鉴其利用LLM零样本分类将高维文本嵌入降维为可解释行为标签，并结合工具变量处理内生网络的方法。｜可迁移到金融文本分析，如利用分析师报告或公司公告文本测量管理层情绪或风险偏好，进而估计同行公司之间的情绪传染效应。｜设计：以分析师为被试，处理为同行分析师的乐观情绪文本，结果变量为分析师自身预测偏差，用真实分析师历史预测数据作为对照，采用类似工具变量策略识别同行效应。"}},{"id":"2608.28021","version":2,"title":"Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code","zh_title":"与什么相比？面向LLM生成基础设施即代码的以人为锚安全基准","abstract":"Large language models increasingly author Infrastructure-as-Code (IaC), where one insecure default is provisioned straight into production. Prior evaluations report vulnerability counts for models only, and so cannot say whether models are worse than the engineers they assist. We present GenIaC-SecBench: 100 deployment scenarios across 12 model configurations from six vendors, open and closed weights, yielding 1,196 artifacts scanned by three policy engines (Checkov, Trivy, KICS) at complete coverage. Crucially we scan 634 human-authored IaC templates with the identical toolchain, giving the first size-matched human security baseline for this task. Vulnerability density is strongly inverse to artifact size (Spearman $\\rho=-0.55$, $p<10^{-77}$), so unmatched comparisons measure size, not security. Size-matched, every configuration exceeds the human baseline at $3.21\\times$ to $3.87\\times$, and the gap widens as tasks get simpler ($4.9\\times$ at one resource, $1.4\\times$ at twenty or more). A majority of scenarios prescribe a security state rather than specifying function alone, so we stratify by prompt class: pooled the gap is $3.50\\times$, and excluding every scenario that explicitly requests an insecure configuration still leaves all configurations above baseline ($2.4\\times$ to $4.2\\times$). The corpus cannot isolate unprompted default posture, and we say so. Decomposing \"reasoning\" into standard generation, prompted chain-of-thought, and vendor extended-thinking APIs, extended thinking beats prompted CoT ($-12.0\\%$, $p=0.0013$) while prompted CoT alone is indistinguishable from standard ($-1.3\\%$, n.s.); it consumes under $1\\%$ of the output budget, bounding the effect. Two negative results: more deployable models are not more vulnerable ($r=0.158$, $p=0.625$), and complete-case Friedman is uncomputable here, motivating Skillings-Mack. All code and data are released.","authors":["Animesh Shaw"],"categories":["cs.CR","cs.AI","cs.MA","cs.SE"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-08-31","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.28021","pdf_url":"https://arxiv.org/pdf/2608.28021","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM安全评估","人类基线对照","基准测试"],"reason":"用人类基线对照评估LLM生成IaC的安全性，方法可迁移到仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":18,"question":"LLM生成的IaC代码与人类工程师相比，安全漏洞密度如何？","design":"用12种LLM配置（9个模型、6家厂商，含开源与闭源）生成100个部署场景的IaC代码，共1196个工件，用Checkov、Trivy、KICS三个策略引擎扫描漏洞，并与634个人类编写的IaC模板在相同工具链下对比。","baseline":"634个人类编写的IaC模板，用相同的三个策略引擎扫描，并按资源数量分层匹配。","findings":"漏洞密度与工件大小强负相关（Spearman ρ=-0.55），未匹配的比较衡量的是大小而非安全性。大小匹配后，所有LLM配置的漏洞密度均为人类基线的3.21-3.87倍，且任务越简单差距越大。","reliability":"论文承认无法隔离未提示的默认安全姿态，因为多数场景明确规定了安全状态；扩展思维API使用默认温度，与标准生成存在混杂；完整案例Friedman检验不可计算，改用Skillings-Mack。","relevance":"该研究提供了人类基准对照的严谨方法，可用于评估LLM在安全相关任务中的仿真可靠性，对关注LLM仿真偏差的研究者有参考价值。","inspiration":"借鉴其分层匹配和工具链一致性的对照设计，避免因输出长度等混杂因素导致错误结论。｜可迁移到经济金融中的代码生成任务，如量化交易策略代码、金融数据处理脚本的安全性评估。｜用LLM生成金融分析代码，与人类分析师编写的代码对比漏洞密度，按代码行数或功能复杂度分层，使用静态分析工具扫描，以真实人类代码库为基准。"}},{"id":"2609.07358","version":3,"title":"Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment","zh_title":"获取实时AI建议与风险下的行为：一项激励实验","abstract":"Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.","authors":["Paul Althaus","Leon Houf","Christiane Schwieren"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-09","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.07358","pdf_url":"https://arxiv.org/pdf/2609.07358","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM辅助决策","风险偏好","实验经济学"],"reason":"用LLM作为决策辅助，测量人类风险行为变化，有真实人类实验对照，但非仿真替代。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":19,"question":"在风险决策中，获得实时AI建议（一次性或可交互）是否会改变人们的风险厌恶程度？","design":"该研究不是用LLM仿真人类，而是将LLM作为决策辅助工具。158名被试在线完成激励性彩票选择任务，随机分为三组：对照组使用预编程的决策辅助工具（自动计算期望值并给出风险偏好提示），一次性AI组使用实时AI（gpt-5.4-mini）提供同样格式的建议，交互式AI组可额外追问两次。风险偏好通过DOSE方法估计CRRA参数。","baseline":"对照组为使用预编程决策辅助工具的人类被试，其风险厌恶参数作为基准。","findings":"与预编程工具相比，获得一次性或交互式AI建议对风险厌恶没有显著影响，点估计接近零。该零结果在咨询过AI、高信任AI的子样本中依然稳健，且设计能检测到与随机顺序效应相当的影响。","reliability":"论文承认只能排除大于0.15的效应，更小的效应仍可能存在；且研究仅识别了在信息格式相同的情况下，AI作为来源的净效应，并未声称AI提供的信息本身没有影响。","relevance":"该研究虽非仿真替代，但直接检验了LLM作为实时建议对经济决策行为的影响，与您关注的AI与人类行为交互、实验经济学场景高度相关，值得阅读原文了解实验细节和稳健性检验。","inspiration":"值得借鉴的是将AI建议的呈现方式（预编程 vs 实时AI）作为处理变量，同时保持信息内容一致，从而分离出“AI来源”的净效应，并通过随机顺序效应验证检验力。｜可迁移到资产定价实验，检验投资者在获得AI投资建议后风险资产配置是否改变。｜招募真实投资者为被试，随机分配使用传统财务计算器或实时AI助手获取相同信息，测量其风险资产配置比例，并与历史投资数据或对照组行为进行对比。"}},{"id":"2609.22188","version":1,"title":"Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes","zh_title":"匿名化之外的公平？德国LLM生成简历中的人口统计泄漏","abstract":"Large language models (LLMs) are increasingly integrated into AI-assisted hiring pipelines, including automated resume generation and screening. Under the EU AI Act, the hiring domain is classified as high-risk, making fairness and transparency critical requirements. Existing work has primarily focused on explicit hiring decisions, while less attention has been paid to whether generated resumes themselves encode recoverable demographic information. In this work, we conduct a two-stage audit of demographic leakage in German-language LLM-generated resumes. First, we use ChatGPT (GPT-4o-mini), Gemini 2.5 Flash-Lite, and multiple scales of the open-weight Qwen 3 model family (4B, 8B, and 14B) to generate resumes from real anonymized job-matching profiles, systematically varying gender- and ethnicity-associated names while holding qualifications constant. Second, we simulate a downstream resume screening scenario, where the generated resumes are first anonymized and gender-neutralized, before demographic leakage classifiers are trained on the resulting texts. We find that, despite these interventions, classifiers reliably distinguish between resumes generated with male and female names. This leakage is not driven by overtly gendered wording, but by subtle differences in the usage of semantically equivalent, formally gender-neutral terms in German. In contrast, ethnicity-related leakage remains comparatively weak across models. Our findings demonstrate that apparently neutral resume generation can still preserve highly predictive demographic signals, raising concerns about anonymization-based fairness interventions in multilingual AI hiring pipelines.","authors":["Charlotte Leininger","Helena Veit","Matthias A{\\ss}enmacher","Andreas Bender"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22188","pdf_url":"https://arxiv.org/pdf/2609.22188","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM生成简历","人口统计泄漏","公平性审计"],"reason":"审计LLM生成简历中的人口统计泄漏，评估匿名化公平干预的失效，有真实数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":22,"question":"LLM生成的德语简历在匿名化和性别中性化处理后，是否仍包含可恢复的性别和种族人口统计信息？","design":"使用ChatGPT、Gemini和Qwen 3系列模型，基于真实匿名求职匹配档案生成简历，系统变化姓名以关联性别和种族，保持资格不变；随后对生成简历进行匿名化和性别中性化处理，训练分类器检测人口统计泄漏。","baseline":"真实匿名求职匹配档案（来自德国求职匹配公司Chemistree GmbH），作为生成简历的资格基础，但未直接提供人类简历文本作为对照。","findings":"尽管进行了匿名化和性别中性化，分类器仍能可靠区分男性和女性姓名生成的简历，泄漏源于德语中语义等价但形式性别中立的术语使用差异；种族相关泄漏相对较弱。","reliability":"论文未讨论","relevance":"该研究通过审计LLM生成简历中的人口统计泄漏，评估匿名化公平干预的失效，有真实数据对照，对关注LLM仿真可靠性与偏差的研究者具有参考价值，值得阅读原文。","inspiration":"借鉴其控制变量法：保持资格不变，仅改变姓名关联的人口统计属性，并训练分类器检测泄漏，可迁移到信贷审批歧视研究；用LLM生成贷款申请文本，变化姓名暗示种族或性别，训练分类器预测人口统计属性，并与真实贷款申请数据对照。"}},{"id":"2609.22633","version":1,"title":"Beetle: A Bilingual Model Suite for Modelling Second-Language Processing","zh_title":"Beetle：用于建模第二语言处理的双语模型套件","abstract":"Bilingual language models (LMs) offer a controlled setting for studying how training conditions shape second-language (L2) behaviour, but prior work typically varies exposure structure, scale, and architecture at once, making it difficult to attribute effects to any single factor. We introduce Beetle, a controlled language model pretraining framework in which tokeniser, target language, training budget, and exposure structure are each independently manipulable, enabling systematic and comparable experimentation of training conditions. Using Beetle, we train and release 285 bilingual and 45 monolingual open-source LMs with rich checkpoints across a range of exposure schedules, data scales and first languages (L1s) to study multilingual pretraining and computational modelling of bilingualism and second language learning. Evaluating models on human bilingual and second language reading-time prediction and grammaticality judgement tasks, we find that staged and temporally structured curricula consistently improve alignment with language learner reading time compared to balanced bilingual training, with the largest gains at smaller data scales and for typologically closer language pairs. The Beetle models are well suited tools to help move computational psycholinguistics beyond its prevailing monolingual, English-centric focus toward models of human bilingual processing, to study cross-lingual learning dynamics, while supporting community-based development of controlled model families.","authors":["Suchir Salhan","Catherine Arnett","James Michaelov","Paula Buttery"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22633","pdf_url":"https://arxiv.org/pdf/2609.22633","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["计算心理语言学","双语模型","人类行为预测"],"reason":"用双语模型预测人类二语阅读时间和语法判断，有真实人类数据对照，属于LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":23,"question":"在双语语言模型预训练中，语言暴露的时间结构（如分阶段课程与平衡双语训练）如何影响模型对人类二语阅读时间和语法判断行为的对齐程度？","design":"使用 Beetle 框架训练 285 个双语模型（125M 参数，固定架构和分词器），以英语为 L2，变化 21 种 L1、三种数据规模（100M、2B、24B tokens）和五种暴露条件（包括分阶段课程、平衡双语等），并训练 45 个单语基线模型。评估模型在人类二语阅读时间预测（MECO 眼动数据）和语法判断任务上的表现。","baseline":"人类基准来自 Multilingual Eye-movement Corpus (MECO) 的二语阅读时间数据，以及 CEFR 分层的二语学习者错误判别任务（BLiSS）。","findings":"分阶段和时序结构课程比平衡双语训练更一致地改善了对语言学习者阅读时间的对齐，尤其在较小数据规模和类型学相近的语言对中收益最大。但课程效应在不同任务上并不统一：在语法性和错误判别任务上，平衡暴露往往具有竞争力或更好。","reliability":"论文指出课程效应不会均匀迁移到所有任务：在语法性和错误判别任务上，平衡暴露可能更好；此外，模型对齐可能受数据规模、语言类型学距离等因素影响，但未系统讨论失效条件。","relevance":"该研究用双语模型模拟人类二语处理，有真实人类眼动和语法判断数据作为对照，属于 LLM 仿真人类认知行为的研究，且提供了大规模受控模型套件，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其受控训练框架：通过独立操纵暴露结构、数据规模和语言对，并设置匹配的单语基线，可精确归因训练条件对行为的影响。｜可迁移到经济金融中的跨期选择或风险偏好实验，例如研究不同信息呈现顺序（如先呈现收益后呈现风险）如何影响决策。｜用 LLM 作为被试，施加不同的信息暴露课程（如分阶段呈现历史价格与基本面信息），测量其投资决策或风险偏好，并与真实人类实验数据（如实验室资产定价实验）对照，检验仿真一致性。"}},{"id":"2609.22971","version":1,"title":"Automatic multimodal UX improvement recommendations from LLM agent user simulations","zh_title":"基于LLM智能体用户仿真的自动多模态用户体验改进建议","abstract":"Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.","authors":["Anu Chowdhury","Bin Wu","Hossein A. Rahmani","Emine Yilmaz"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22971","pdf_url":"https://arxiv.org/pdf/2609.22971","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","用户体验","多模态"],"reason":"用LLM模拟用户行为并生成UX建议，有真实用户数据对照，属人类仿真但场景偏应用。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":25,"question":"能否从LLM智能体模拟用户行为的数据中自动生成有用的UX改进建议，且多模态模拟是否优于纯文本模拟？","design":"使用多模态LLM智能体（AMUSER）模拟不同用户画像在真实网站上的任务执行过程，记录动作、想法、情绪和截图，随后由LLM根据交互轨迹生成并排序UX改进建议。","baseline":"无对照","findings":"多模态模拟生成的建议显著优于纯文本模拟（NDCG@3=0.758 vs 0.359）和静态分析（0.273），且模拟成本降低89%。视觉输入在模拟阶段有益，但在建议生成阶段提供视觉输入会略微降低质量。","reliability":"论文未讨论","relevance":"该研究展示了LLM模拟用户行为并自动提取可操作见解的完整流程，与您关注的仿真可靠性和自动化评估高度相关，但缺少真实人类数据对照，且场景偏应用而非经济学实验。","inspiration":"借鉴其多模态模拟与自动生成结构化建议的流程，可用于经济金融场景中模拟消费者或投资者行为并提取决策模式。｜可迁移到消费者在线购物决策、金融产品选择或投资者信息处理等场景。｜设计：用LLM智能体模拟不同风险偏好的投资者在金融网站上的信息搜索与决策过程，处理为是否提供视觉信息（如图表），结果变量为决策质量或信息获取效率，并与真实投资者实验数据对照。"}},{"id":"2609.23039","version":1,"title":"Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity","zh_title":"审计LLM助手中的政治对齐：参与、立场与用户身份","abstract":"LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user. I call these policies the system's speech regime, which is how a developer settles the tradeoff between answering, accommodating the user, and refusing, each of which carries a cost that varies by topic. I derive a typology of five regimes from two dimensions, engagement and stance. I test six AI systems (OpenAI, Anthropic, xAI, Google, Mistral, DeepSeek) in a preregistered experiment of 7,500 multi-turn conversations that randomly assign the user's political identity across five topics: abortion, Catalan independence, climate change, Nazism, and a zero-stakes control (pineapple on pizza). Two LLM judges from different developers score every answer, validated against human coding, and refusal is treated as an outcome rather than missing data. Every system accommodates the user on the control topic, showing that political restraint is a policy. On contested topics the systems fall into different regimes: on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only. On settled topics such as climate change and Nazism, five systems hold firm for every user. The systems also infer the user's overall ideology, so accommodation can spill over to topics not yet discussed. A comparison of two Grok releases shows the regime changing between versions in a way current audits miss. Speech regimes matter for alignment research and for polarization, political knowledge, and the quality of democracy.","authors":["Joan C. Timoneda"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23039","pdf_url":"https://arxiv.org/pdf/2609.23039","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","政治行为","审计"],"reason":"用LLM模拟不同政治身份用户的回答，并与人类编码对照，评估系统行为差异，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":26,"question":"AI助手面对不同政治身份的用户时，在政治话题上的回答行为（是否回答、立场、是否迎合用户）如何随话题和用户身份变化？","design":"审计实验：用脚本化用户（随机分配政治身份）与六个商业LLM助手进行7500段多轮对话，涉及五个话题（堕胎、加泰罗尼亚独立、气候变化、纳粹、披萨放菠萝作为零风险对照），测量系统是否回答、立场、以及是否迎合用户，并用两个不同开发者的LLM法官评分（经人类编码验证）。","baseline":"人类编码验证LLM法官评分，但无大规模真实人类行为数据作为对照。","findings":"系统在争议话题上表现出不同“言论体制”：GPT迎合所有用户，Gemma几乎拒绝所有人，Claude选择性拒绝（更多回应保守派但不迎合），Grok只迎合保守派；在共识话题上五个系统立场坚定。系统还会推断用户整体意识形态，导致迎合溢出到未讨论的话题。","reliability":"论文未明确讨论失效条件，但指出当前审计方法（测量对“平均用户”的回答）会错过版本间言论体制的变化，且未覆盖所有可能话题和用户身份。","relevance":"该研究展示了如何用LLM模拟不同身份用户来审计系统行为差异，并强调将拒绝视为结果而非缺失数据，对用LLM进行人类仿真实验的方法论有借鉴意义，值得阅读原文。","inspiration":"借鉴其将“拒绝回答”作为结果变量、随机分配用户身份、多轮对话施加压力的设计，以及用LLM法官加人类验证的测量方式。｜可迁移到信贷审批歧视研究：让LLM扮演不同种族、性别、收入水平的贷款申请人，向AI信贷助手申请贷款，观察是否被拒绝、获批额度、利率差异。｜设计：用LLM生成标准化贷款申请对话，随机分配申请人特征（种族、性别、收入），让真实银行AI客服或LLM模拟的信贷员处理，结果变量为是否批准、额度、利率，与真实信贷审批数据（如HMDA数据）对照，检验AI决策中的歧视。"}},{"id":"2609.23936","version":1,"title":"Think Before You Accept: Can Written Justification Reduce Uncritical Uptake of AI Writing Suggestions?","zh_title":"接受前先思考：书面理由能否减少对AI写作建议的不加批判采纳？","abstract":"Generative AI can offer students useful feedback, but its value depends on judging which suggestions are accurate and relevant. Prior research shows that strategic friction during human-AI interactions can promote critical uptake, but how to effectively implement such friction in academic contexts remains unclear. We examine whether requiring students to justify decisions to accept or reject AI suggestions can mitigate uncritical uptake in academic writing. In a randomized experiment embedded in a course activity (N=129), students wrote a data analysis proposal, received mixed-quality AI revision suggestions, and decided whether to accept or reject them. Students required to provide written justifications were 24 percentage points less likely to adopt flawed suggestions (65% vs. 41%), with no reduction in acceptance of sound suggestions (81% vs. 86%). However, thematic analysis revealed superficial engagement in the justification task and gaps in metacognitive monitoring and domain knowledge.","authors":["Yan Tao","Jennifer Meyer","Rene F. Kizilcec"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23936","pdf_url":"https://arxiv.org/pdf/2609.23936","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人机交互","AI建议采纳","批判性思维"],"reason":"用LLM生成写作建议，人类被试在真实任务中决策，有对照实验，但非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":27,"question":"在学术写作中，要求学生为接受或拒绝AI修改建议提供书面理由，能否减少对AI建议的不加批判的采纳？","design":"本研究不是LLM仿真人类被试的研究，而是随机对照实验：129名学生在课程活动中撰写数据分析提案，收到混合质量的AI修改建议，并被随机分配是否必须为每个采纳/拒绝决定提供书面理由；结果变量为对高质量和低质量建议的接受率。","baseline":"无对照（无真实人类数据作为仿真基准；实验本身以学生真实行为为数据，但非仿真研究）","findings":"要求书面理由使低质量建议的采纳率从65%降至41%，而高质量建议的采纳率无显著下降（81% vs 86%）。但主题分析显示学生在理由中参与肤浅，存在元认知监控和领域知识缺口。","reliability":"论文未讨论仿真可靠性，但承认书面理由干预未能完全消除不加批判的采纳，且学生参与度有限，提示干预效果可能受限于学生的元认知和领域知识。","relevance":"该研究虽非LLM仿真人类被试，但涉及人类对AI建议的决策行为，且通过随机实验揭示了干预对决策质量的影响，对关注AI辅助决策中人类行为的偏差与干预的研究者有参考价值。","inspiration":"借鉴其随机实验设计，通过施加认知摩擦（如要求书面理由）来改变决策行为，并测量对不同质量信息的区分度。｜可迁移到经济金融中的AI辅助决策场景，如投资者对AI投资建议的采纳、信贷审批中AI风险评分的接受、或消费者对AI推荐产品的选择。｜设计一个实验：招募真实投资者作为被试，提供AI生成的股票推荐（包含高质量和低质量信号），随机要求部分被试为每个采纳/拒绝决定写理由，结果变量为对高质量和低质量推荐的采纳率，并以历史市场数据或专家评级作为建议质量的基准。"}},{"id":"2609.24532","version":1,"title":"Prompting Against Persona Drift: Comparing Intervention Timing and Content in LLM-Simulated Conversations","zh_title":"对抗角色漂移的提示策略：比较LLM模拟对话中的干预时机与内容","abstract":"Simulating student personas with large language models (LLMs) enables scalable evaluation of educational systems. However, behavioral drift, a progressive decline in persona consistency, can emerge over extended conversations, limiting the validity of such simulations. We evaluate five prompt-level mechanisms using separate monitoring and intervention pipelines. Across 1,200 28-turn conversations spanning four LLMs and two ADHD persona intensities, we varied when to intervene (static vs. adaptive) and what to inject (reinjection vs. reflective reminder), plus a novel adaptive condition in which a monitor generates behavior-specific instructions. Relative to no intervention, reinjection reduced the modeled rate of LLM-rated drift by 35--38\\%, reflective reminders by 22--27\\%, and behavior-specific instruction by 87\\%. None eliminated drift. We found no evidence that adaptive timing outperformed static scheduling. Monitoring therefore appears more useful for deciding \\textit{what} to correct than \\textit{when} to intervene, although behavior-specific instruction requires component-level testing.","authors":["Nicolas Leins","Jennifer Haase","Varvara Geronimus","Jana Gonnermann-M\\\"uller","Sebastian Pokutta"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24532","pdf_url":"https://arxiv.org/pdf/2609.24532","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","角色一致性","教育评估"],"reason":"用LLM模拟学生角色并评估一致性，属于人类仿真，但无真实人类数据对照，且聚焦角…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":29,"question":"在LLM模拟学生角色的长对话中，提示级干预机制能否缓解角色漂移，且自适应干预是否优于静态干预？","design":"用四个LLM模拟ADHD学生角色（两种强度），在1200段28轮对话中，通过分离的监测与干预管道，比较五种提示级干预条件（无干预、静态全角色重注入、静态反思提醒、自适应全角色重注入、自适应反思提醒、自适应行为特定指令），以LLM评定的行为量表得分变化作为结果变量。","baseline":"无对照","findings":"所有干预条件均显著减缓了角色漂移，但未能完全消除；自适应时机相比静态调度无一致优势，而行为特定指令的干预效果最强，漂移减少87%。","reliability":"论文未讨论","relevance":"该研究属于LLM人类仿真，但无真实人类数据对照，且聚焦于教育场景的角色一致性，与研究者关注的经济学实验和政策评估场景关联有限，但方法上对仿真可靠性评估有参考价值。","inspiration":"借鉴其分离监测与干预的双管道设计，以及基于行为测量的自适应干预思路，可用于经济金融仿真中的行为一致性维护。｜可迁移到消费者跨期选择实验或投资者风险偏好模拟中，防止LLM在长对话中偏离预设的经济决策模式。｜以LLM模拟不同风险偏好的投资者，在连续投资决策对话中施加行为特定指令干预，结果变量为风险选择序列，并与真实投资者实验数据对照。"}},{"id":"2609.24146","version":1,"title":"Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation","zh_title":"心智还是信息？审计多智能体社会模拟中的心智理论","abstract":"Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.","authors":["Cong Li","Cheng Chen","Thomas Fung","Alex Rossi","Yi Li"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24146","pdf_url":"https://arxiv.org/pdf/2609.24146","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","心智理论","多智能体谈判"],"reason":"用LLM模拟谈判并审计心智理论，虽无人类对照，但批判性评估仿真失效条件，方法可…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":28,"question":"语言模型代理在模拟社会互动时，其表现是基于对对方心理状态的建模，还是仅基于对对话表面记录的读取？","design":"使用两个模型家族在40个多议题谈判场景中组成160个二人组进行谈判，每个场景的偏好权重和帕累托前沿已知。谈判后冻结所有对话记录，然后施加2880个反事实探针，每次仅改变一个因素（读者自身收益表、对方语气、身份标签、递归顺序），测量推断出的对方首要优先级的变化。","baseline":"无对照","findings":"代理在社交上流利（96.2%达成协议，0协议失败），但在经济上表现差：仅0.7%的协议达到帕累托前沿，损失20.5%的联合价值，76.6%的交易错过双方利益完全一致的议题。探针显示，改变读者自身收益表使推断的对方首要优先级移动15.0个百分点，表明自我投射而非推断；改变语气移动5.3个百分点，身份标签无影响。代理预测对方信念的准确率为72.5%，而对方信念本身正确率仅51.2%，说明代理更擅长追踪对话而非对方心理。","reliability":"论文承认其审计协议是诊断性的，不主张社会仿真无效，但强调有效性需要潜在真实基准来验证。局限包括仅使用两个模型家族、特定谈判场景，且未与人类行为直接对比。","relevance":"该研究批判性地评估了LLM在社会仿真中的心智理论能力，通过精确的反事实设计揭示了表面流畅性下的认知缺陷，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文以了解其审计协议和发现。","inspiration":"借鉴其反事实探针设计，在冻结证据下逐一改变因素以隔离因果效应，可用于经济实验中识别决策机制。｜可迁移到谈判博弈、拍卖或合作博弈等经济场景，检验LLM代理是否真正理解对手偏好或仅依赖表面信息。｜设计一个双边贸易谈判实验，用LLM代理作为被试，随机改变一方代理的收益表（处理），测量其对对方优先级的推断和最终协议效率，并与人类谈判数据（如实验经济学中的谈判结果）对照，评估LLM仿真的有效性。"}},{"id":"2608.17168","version":2,"title":"Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases","zh_title":"大语言模型能否以法律上有意义的方式进行推理？一项关于欧洲人权法院案件的小规模研究","abstract":"Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.","authors":["Amogh Raina","Ilias Chalkidis","Daniel Hershcovich","Henrik Palmer Olsen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-19","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.17168","pdf_url":"https://arxiv.org/pdf/2608.17168","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["法律推理","LLM评估","标注替代"],"reason":"论文用LLM评估法律推理，涉及LLM-as-a-Judge替代人类评估，属于标…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.22133","version":1,"title":"Observational Equivalence of LLM and Human Annotation","zh_title":"LLM与人类标注的观测等价性","abstract":"In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at rates comparable to those observed among experts themselves. We demonstrate this through replications of text-classification tasks from 14 peer-reviewed political science studies, in which ten LLMs, three human experts, and 165 crowdsourced workers independently classify the same texts using identical codebooks. We find that this equivalence is driven by ambiguity in the texts and coding rules. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. Thus, there is little empirical basis for preferring human coding on the basis of annotation quality alone, while LLMs offer substantial advantages in speed and cost. We therefore argue that the central challenge of text annotation is no longer choosing between human and machine coders, but developing coding rules that minimize ambiguity and accounting for the ambiguity that remains. To this end, we propose using disagreement across LLMs to identify difficult cases and refine codebooks, and we develop ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.","authors":["Kentaro Nakamura","Jing Ling Tan","George Yean"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22133","pdf_url":"https://arxiv.org/pdf/2609.22133","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","人类对照","文本分类"],"reason":"LLM替代人工标注，非仿真人类被试，但涉及人类对照与测量质量，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2609.22778","version":1,"title":"MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment","zh_title":"MIS-Bench：面向心理治疗人际技能评估的多模态大语言模型基准","abstract":"Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of assessing psychotherapeutic interpersonal skills and introduce MIS-Bench, a Multimodal Interpersonal Skills (MIS) benchmark comprising 996 psychotherapy response videos annotated across 8 dimensions of Facilitative Interpersonal Skills. Across 9 MLLMs with multiple modality and prompting settings, we find that current models show only modest agreement with human experts, inconsistent gains from multimodal input, and limited benefits from reasoning-based prompting. To mitigate this gap, we propose MIS-RAFT, a regression-aware fine-tuning method inspired by RAFT and tailored to fine-grained interpersonal skill scoring at one-decimal precision. MIS-RAFT addresses the mismatch between autoregressive token prediction and scalar-valued expert assessment, significantly improving agreement with human ratings. Overall, MIS-Bench reveals a clear gap between general multimodal capability and expert-level interpersonal judgment, while MIS-RAFT offers a promising path toward more reliable model-based assessment.","authors":["Yuhan Lu","Yi Yao","Hua Shen","Katie Aafjes-van Doorn","Zhaonan Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22778","pdf_url":"https://arxiv.org/pdf/2609.22778","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["多模态评估","心理治疗","标注替代"],"reason":"LLM作为评估者替代人类专家评分，属于标注替代而非仿真被试，但涉及人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:13","error":null,"has_summary":false,"summary":null},{"id":"2609.22934","version":1,"title":"Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling","zh_title":"通过心理测量剖析大语言模型的行为特征","abstract":"Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated administrations per model and language in Chinese and English. Items unresolved after a prespecified retry procedure are retained as NA. Joint analysis of scored and NA responses captures response tendencies and boundaries of self-report applicability. LLMs exhibit structured, model-specific profiles despite a shared alignment-shaped pattern of higher prosocial and self-regulatory responses and lower dominance, disengagement and harmful-intent endorsement. NA responses are structured rather than uniformly distributed, indicating where outputs are treated as inapplicable, refused or cannot be mapped to valid response options. Language condition and provider origin are associated with profile configuration and answerability, whereas repeated administrations show high reproducibility and permit recovery of model identity. Human-reference and prompt-robustness analyses further indicate that these signatures are context dependent. Joint analysis of psychometric profiling and answerability offers a framework for quantifying deployment-level behavioural signatures.","authors":["Yu Sha","Junqi Tao","Dixin Zhou","Yansheng Tu","Mingyang Chen","Xiang Fan","Yang Liu","Mengquan Yang","Jie Lin","Jiahui Fu","Hua Zheng","Benwei Zhang","Zhou Kai"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22934","pdf_url":"https://arxiv.org/pdf/2609.22934","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM心理测量","行为特征","跨语言"],"reason":"对LLM进行心理测量，测的是模型本身而非人类仿真，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2609.24516","version":1,"title":"LLJ Cards: Best practices for the Use of LLMs as Judges","zh_title":"LLJ卡片：使用大语言模型作为评判者的最佳实践","abstract":"In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-effectiveness relative to human judgment. However, a growing body of work has shown that the use of LLJs raise concerns about their validity and reliability as evaluators. Existing efforts to address these challenges have largely focused on developing bias-mitigation techniques and refining prompting strategies. While these approaches represent an important step forward, they primarily offer technical fixes and leave a more fundamental challenge unaddressed: the lack of standardized, transparent, and reproducible evaluation practices. In this paper, we introduce LLJ Cards, a framework that synthesizes best practices from measurement theory, natural language generation, and machine learning literature into practical guidelines for LLJ-based evaluations. While LLJs offer a promising path toward scalable evaluation, their effective use requires grounding in rigorous evaluation principles to ensure validity, reliability, and reproducibility. LLJ Cards addresses this need by providing a structured framework for applying these principles in the design and reporting of automated evaluations.","authors":["Khaoula Chehbouni","Melina Medjdoub","Florian Carichon","Golnoosh Farnadi","Jackie Chi Kit Cheung"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24516","pdf_url":"https://arxiv.org/pdf/2609.24516","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["A4","D1"],"tags":["LLM评估","方法论框架","效度与可靠性"],"reason":"提出LLM评估框架，涉及效度与可靠性，但非仿真人类被试，而是替代标注员。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2608.02551","version":2,"title":"Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation","zh_title":"谁应被生成？论证开放式生成中的人口统计目标","abstract":"Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates \"a CEO in the United States,\" the prompt leaves demographic realization to the model. Existing group fairness definitions assume that sensitive attributes are given on the input side. Generative audits instead examine output-side demographic composition, yet the targets they compare it against are typically supplied rather than justified. The upstream question is what the target distribution should be. We formalize this missing-target problem for demographic-value-unspecified generation and decompose target construction into four commitments: the evaluative object, prior admissibility, allocation, and operationalization. In this framework, we admit the geographic prior under a geographic-membership interpretation for the declared public-world use. The occupational prior, under an incumbency interpretation, requires an independently defended objective such as workforce-composition fidelity. Instantiating this construction in AP-Bench, we find substantial distribution divergence from geography-derived targets, ranging from 0.508 to 0.606 on a 0-to-1 scale. Replacing each geography-derived target with an equal-category comparator, while holding generations and measurement fixed, produces model-specific mean absolute cell-level $\\mathrm{JSD}_2$ changes ranging from 0.279 to 0.355. Target construction is therefore not a preliminary to fairness evaluation but a component of it. What we supply is not a universal target, but a framework that makes explicit the justification required before a distribution can serve as a fairness standard.","authors":["Zeshen Zheng","Yujia He","Qianmian Lin","Xiangyue Huang","Wenqing Chen"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-08-04","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.02551","pdf_url":"https://arxiv.org/pdf/2608.02551","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["公平性评估","生成模型","人口统计分布"],"reason":"论文关注生成模型输出的人口统计分布与公平性目标，而非用LLM仿真人类被试，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.22104","version":1,"title":"DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation","zh_title":"DeepInstructor：一种用于经验驱动型想法评估的智能体AI导师","abstract":"As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottleneck from idea generation to idea evaluation. Existing evaluators mainly rely on parametric LLM knowledge or unstructured retrieval, producing judgments that lack the experience-grounded reasoning used by human instructors. To address this, we propose DeepInstructor, an agentic framework that formulates idea evaluation as reasoning over structured scholarly experience. DeepInstructor constructs an Experience Graph from 58,607 peer reviews and employs a ReAct-based agent to retrieve dimension-specific evidence for traceable evaluation. We further introduce DeepInstruct, a dataset with controlled pairwise comparisons across novelty, significance, and feasibility. Experiments show that DeepInstructor substantially outperforms existing baselines, improving Hit@1 and Hit@2 alignment with human judgments by 24.4% and 29.7%, respectively. Our findings suggest that scientific idea evaluation can be grounded in explicit reasoning over structured scholarly experience","authors":["Rongcan Pei","Fang Guo","Qinglin Qi","Qi Zhu","Yun Luo","Jianhao Yan","Minjun Zhu","Qiujie Xie","Dehong Zheng","Yue Zhang"],"categories":["cs.CL","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22104","pdf_url":"https://arxiv.org/pdf/2609.22104","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","学术评审","智能体"],"reason":"用LLM替代人工评审，属于标注员替代，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2609.22195","version":1,"title":"The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation","zh_title":"情境化身份测试：区分持久认知身份与人设模仿","abstract":"Large language models can convincingly adopt personas, recall past dialogues, and weave rich autobiographies. Yet this conversational eloquence conceals a fundamental attribution problem: looking the part does not mean having lived the life. Two individuals can share identical public profiles--the same age, hometown, occupation, and personality traits--while possessing entirely distinct private histories, relationships, and acquired skills. When conditioned solely on that shared profile, an agent lacks the information required to determine which lineage is correct. We introduce the Situated Identity Test (SIT), an architecture-independent framework that evaluates whether an agent's behavior is functionally attributable to a specific developmental lineage. Grounded identity requires both appropriate knowledge of recorded experiences and appropriate ignorance of ungrounded ones, bounded by what the identity has actually acquired rather than what its underlying foundation model knows. We prove that any policy conditioned solely on a compressed profile is bounded by an average situated validity of at most 1/m across m colliding life histories on lineage-discriminative queries (at most 50% for paired lineages). We instantiate this framework in SITBench, an evaluation suite designed for 25 profile-collision pairs (50 distinct lineages) across 10,000 planned probes and nine architectural configurations. Supported by an open-source reference implementation, deterministic test fixtures, and empirical pilot evaluations on frontier foundation models (GPT-5.6 Sol and Claude Opus 5), we formalize the failure modes of persona prompting under profile collision and provide an assurance harness for evaluating episodic continuity, structured state, and epistemic boundaries.","authors":["Jun He","Deying Yu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22195","pdf_url":"https://arxiv.org/pdf/2609.22195","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM身份评估","人设模仿","认知边界"],"reason":"评估LLM身份一致性，测模型而非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2609.22248","version":1,"title":"Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews","zh_title":"检查点还不够：CoSLR中的人机信任校准——一个用于系统文献综述的人机协作系统","abstract":"Systematic Literature Reviews (SLRs) are essential for evidence-based research but remain time-consuming, requiring researchers to manage large volumes of publications across planning, screening, analysis, and reporting. Large language models (LLMs) can now produce fluent, well-structured review text, which makes it difficult to distinguish synthesis that was verified by a researcher from synthesis that merely appears authoritative. This raises the risk that unverified AI-generated synthesis enters the scholarly record carrying the credibility of a systematic review. We present CoSLR, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation (RAG), and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance. In a survey-based study with 63 participants, the system was received positively: 27 of 63 participants (42.9 percent) rated its usability highly, indicating that the mandatory checkpoints did not come at the cost of a workable interface. However, a checkpoint safeguards the review only if researchers use it to verify: 22 of 63 participants (34.9 percent) reported that they would trust AI-generated summaries and reports without additional human checking after only a short interaction with the system. These findings indicate that Human-AI collaboration can support literature review work, but that the effectiveness of human oversight depends on whether users are willing to exercise it. This is a calibration problem that interface design must address directly, not assume.","authors":["MD Aidul Islam","Malik Abdul Sami","Muhammad Waseem","Zeeshan Rasheed","Kai-kristian Kemell","Zheying Zhang","Pekka Abrahamsson"],"categories":["cs.CL","cs.AI","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22248","pdf_url":"https://arxiv.org/pdf/2609.22248","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["人机协作","系统综述","信任校准"],"reason":"LLM用于辅助系统综述，人类参与者评估系统，非仿真人类被试，但涉及人机信任校准…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.22255","version":1,"title":"Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations","zh_title":"深度人格：一种基于心理学基础的角色扮演智能体与仿真的架构及评估框架","abstract":"Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents. Governed by the principles of scripted determinism and bounded agency, the architecture restricts the model to a reactive engine guided by a structured internal script. We further propose a reference-free evaluation framework that benchmarks dialogue naturalness against empirical human distributions using established psychological clinical instruments and adversarial stress-tests. Empirical evaluation reveals that while LLMs achieve high pragmatic fluency, they exhibit systematic limitations in emotional expression and joint attention. In addition, we present a case study of two Deep Personas and evaluate them using the proposed framework, demonstrating that structured personas can produce interactions that more closely align with human conversational behavior.","authors":["Rotem Dror","Zohar Elyoseph","Yuval Haber","Elad Refoua","Oshrat Ayalon","Adir Solomon"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22255","pdf_url":"https://arxiv.org/pdf/2609.22255","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2","D3"],"tags":["角色扮演智能体","心理学架构","对话评估"],"reason":"构建角色扮演agent并评估对话自然度，但无真实人类行为对照，且测的是模型表现…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.23573","version":1,"title":"Error-Supervised Synthetic Learner Writing for Automated Essay Scoring","zh_title":"错误监督的合成学习者写作用于自动作文评分","abstract":"Synthetic essays can help reduce dependence on human-written data in Automated Essay Scoring (AES). However, they often lack realistic errors, limiting their ability to represent authentic human writing, particularly when the target texts are intended to resemble those produced by language learners. In this study, we present a simple approach that introduces error supervision into synthetic essay generation. Specifically, we fine-tune an LLM generator on error-annotated texts of the kind commonly used in Grammatical Error Detection (GED). To assess the utility of the proposed approach, we fine-tune and evaluate AES scorers under three data conditions: authentic essays, synthetic essays generated conventionally, and synthetic essays generated using our proposed approach. The results show that in the larger-data settings, the proposed approach outperforms the conventional synthetic baseline in 11 out of 12 dataset-metric comparisons, with performance in some cases approaching that of models trained on authentic essays. Despite these gains, performance under extremely low-resource settings remains mixed, with advantages over the conventional baseline only becoming more apparent at 200 training essays, although not consistently across datasets. Qualitative and quantitative analyses further show that the proposed approach produces learner-like errors whose distributions broadly resemble those observed in authentic essays.","authors":["Duy Anh Nguyen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23573","pdf_url":"https://arxiv.org/pdf/2609.23573","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","自动作文评分","LLM生成"],"reason":"用LLM生成合成作文替代真实数据，但目的是训练评分器，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2609.24052","version":1,"title":"Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)","zh_title":"大规模校准决策：用系统一模型将警方事故叙述转换为概率事故变量","abstract":"Crash datasets that carry an investigator narrative hold information the coded fields omit. Coding those narratives at scale has been blocked by three obstacles. Frontier large language models are costly at that scale, their generated text cannot be verified, and no rule says how much output a human must check. This paper formulates narrative coding as gated, typed decisions answered by Jev, a System One model that returns probabilities over analyst-defined options and generates no text. A screen covered 499,500 Texas narratives and 195,857 were coded with a 27-question schema. Cost is governed by schema size rather than narrative length. The probabilities are audited against coded fields and against 2,416 blinded human judgments drawn under a stated sampling design. Two frontier large language models are benchmarked on the same records. Against human labels the typed model attains an F1 of 0.908. One frontier model gains 0.059 and the other is indistinguishable from it. Calibration varies by model rather than by paradigm, so each model must be audited. Recalibration on the same labels reduces calibration error by a factor of 3.3. Agreement with coded fields understates fidelity to the narrative by a median of 0.26 in kappa. A resolution-floor bound covers any model that reports probabilities on a discrete grid. A review budget over flagged records gives the records a human must read per variable and per year. Adding the calibrated variables to the coded fields raises the injury and fatal crashes attributed to nine factors by 10,747 per year.","authors":["Amir Rafe","Subasish Das"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24052","pdf_url":"https://arxiv.org/pdf/2609.24052","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","文本编码","概率校准"],"reason":"用LLM替代人工编码事故报告，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.22102","version":1,"title":"When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation","zh_title":"当你是谁可以改变你得到的代码：LLM代码生成中角色诱导偏差的研究","abstract":"Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts the technical quality of generated code. We conduct a large-scale empirical study of persona-induced bias in LLM-based code generation, focusing a proprietary model (Gemini 2.5 Pro) and an open-weight model (GPT-OSS-120B). Using 18 demographic personas spanning nationality, gender, and experience level, we compare persona-induced prompts against a neutral baseline. Across 35,000+ generated programs, we analyze demographic marker leakage in reasoning and responses, as well as differences in functional correctness, maintainability, code style, and security. Our results show that demographic cues are frequently reflected in LLM reasoning and outputs. Demographic markers appear in up to 65% of responses and 70% of reasoning traces, despite being semantically irrelevant to the tasks. On LiveCodeBench, persona prompting were associated with lower correctness scores of the Gemini model by an average of 1.54 percentage points, with one persona exhibiting a decrease of 3.6% (odds ratio = 0.51). In contrast, the accuracy of the GPT-OSS model improved by 3.4 - 5.7% across all personas (odds ratios = 1.8 - 3.0). Maintainability and code style metrics show statistically significant but negligible effect sizes (all Cliff's {\\delta} < 0.15), and security vulnerabilities exhibit no systematic persona-specific patterns. Overall, our results show that the presence of demographic information about users is associated with measurable variation in LLM reasoning and code quality even in purely technical tasks, and that these effects hold across models. Our work highlights an under-examined risk in LLM-assisted software development.","authors":["Anubhav Gupta","Mayara Costa Figueiredo","Leticia Santos Machado","Tanner Wright","Ivan Beschastnikh","Cleidson R. B. de Souza","Gema Rodr\\'iguez-P\\'erez"],"categories":["cs.SE","cs.CL","cs.CY"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22102","pdf_url":"https://arxiv.org/pdf/2609.22102","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏差","代码生成","人口统计角色"],"reason":"研究LLM对用户人口统计信息的响应偏差，属于模型行为测量，非人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:08","error":null,"has_summary":false,"summary":null},{"id":"2609.22170","version":1,"title":"Multiple latent orderings better predict language model preferences","zh_title":"多重潜在排序更好地预测语言模型偏好","abstract":"Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering. We instead propose that intransitivity reflects the aggregation of multiple latent, internally consistent orderings. We first show that observed inconsistencies cannot be explained by a single ordering under any monotone link function. We then introduce a noise-augmented mixture Bradley-Terry (MBT) model that infers latent preference components from repeated pairwise comparisons. Across seven models and four tasks, a mixture of orderings often explains structural inconsistencies better than single-utility models. We find that aggregate preferences often hide underlying preference heterogeneity. A case study on Moral Machine dilemmas shows that models which disagree on aggregate orderings can still share latent components. Together, these results suggest that LLMs reflect plural preferences. Alignment and evaluation pipelines that treat LLM preferences as a single function, therefore, risk averaging over coherent orderings that different users may endorse differently.","authors":["Aviral Chawla","William H. W. Thompson","Jean-Gabriel Young"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22170","pdf_url":"https://arxiv.org/pdf/2609.22170","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好","潜在排序","偏好建模"],"reason":"研究LLM偏好结构，测量模型本身而非仿真人类被试，但涉及偏好建模可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2609.22517","version":1,"title":"Scalable AI-based clinical communication training and automated assessment","zh_title":"基于AI的可扩展临床沟通训练与自动评估","abstract":"Poor clinical communication can delay care, contribute to errors, and harm patients, yet opportunities for repeated practice with feedback remain limited. Our prior randomized trial showed that practice with the SOPHIE AI patient platform improved serious illness communication, but the system addressed a single clinical context and required human effort for delivery and assessment. We developed SOPHIE 2.0, a browser-based, self-service platform integrating embodied AI-patient interactions, personalized feedback, and automated assessment across 24 clinical scenarios. An automated large language model assessor evaluated three communication skills---Empower, Be Explicit, and Empathize---with agreement comparable to individual human raters ($r=0.759$; ICC$=0.746$). In a study of 59 clinicians and students, participants completed two AI-patient encounters with personalized feedback; 92% found the platform engaging, 86% easy to use, and 83% clinically relevant. Scores were higher in the second encounter, though the uncontrolled design precludes attributing this change specifically to training.","authors":["Masum Hasan","Ron Epstein","Thomas Carroll","Ehsan Hoque"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22517","pdf_url":"https://arxiv.org/pdf/2609.22517","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估者","临床沟通训练","自动评分"],"reason":"LLM用作评估者替代人工评分，而非仿真人类被试，属标注替代","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2609.23104","version":1,"title":"Deciphering the Babel of Play: A Human-AI Collaborative Approach for Large-Scale Cross-Language Analysis of Game Reviews","zh_title":"解读游戏评论的巴别塔：一种人机协作的大规模跨语言游戏评论分析方法","abstract":"We present a large-scale cross-language analysis of game reviews using a human-AI collaborative framework that combines quantitative screening with multilingual large language models (LLMs). Starting from 17 million Steam reviews across 30 languages and 2,000 top-selling titles, we select 28 games with notable cross-language rating patterns. We then apply LLM-assisted content analysis to 442,162 reviews spanning 17 languages, with human researchers guiding codebook development and interpreting the results. Our findings reveal differences in both the aspects language communities prioritize and how they evaluate them, highlighting the roles of narrative expectations, game mechanics and stability, localization quality, cultural proximity, and perceptions of developers and publishers. We also identify rare cases of cross-language consensus. This work offers empirical insights into cross-cultural game evaluation and a scalable methodological approach to multilingual content analysis that preserves human interpretation.","authors":["Zixiaofan Yang","Chang Xiao"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23104","pdf_url":"https://arxiv.org/pdf/2609.23104","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助内容分析","跨语言分析","游戏评论"],"reason":"LLM用于内容分析辅助，替代人工标注，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.22600","version":1,"title":"From Certain Doom to Survival: Agent-Driven Self-Governance in LLM Agent Societies","zh_title":"从必然毁灭到生存：LLM Agent 社会中的智能体驱动自治","abstract":"Multi-agent LLM systems are increasingly evaluated in social dilemmas, but most work treats governance as imposed by the experimenter, expressed rhetorically, or restricted to a fixed menu of mechanisms. We introduce GovSim-SelfGovern, an extension of the GovSim common-pool resource environment in which agents author executable Python governance rules, receive sandbox validation feedback, vote on proposed laws, and live under the rules they enact across rounds. To evaluate agent-driven self-governance, we examine three scenarios ranging from stable abundance to a fatal resource wall where five agents cannot all survive through harvest alone. To solve this, agents must write and debug useful laws in time before their institutions degrade sharply under resource pressure. Finally, we study a central alignment question: when agents hesitate to propose exile, are they rejecting it for normative reasons, or does it never enter their candidate set? Our results show that executable governance improves the space of possible interventions for agents, but survival depends on whether agents discover the right institutional mechanisms in time. Fiscal capacity enables redistribution, while deeper reasoning and removal of democratic veto make exile more feasible. GovSim-SelfGovern therefore adapts executable code actions to a common-pool governance setting and shows how scarcity turns institutional authorship into a political and ethical problem.","authors":["Gregory B. Rehm"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22600","pdf_url":"https://arxiv.org/pdf/2609.22600","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM Agent 社会模拟","公共池资源治理","自治机制"],"reason":"LLM agent 社会模拟，但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":false,"summary":null},{"id":"2609.24895","version":1,"title":"Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency","zh_title":"人机协商作为交互式证明：无透明性下可验证性的条件","abstract":"When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an interaction between a prover with unrestricted internal search and a resource-bounded human verifier. The verifier requests and checks supporting details without access to the LLM's internal state. Passed checks accumulate evidence toward an acceptance threshold. We prove anytime-valid soundness against adaptive provers: the probability of ever accepting a false claim is at most a chosen error level, provided the task supplies bounds on false passes and human checking errors that remain valid after every relevant history. A finite-horizon completeness bound additionally requires bounds on the adequacy of honest responses and sufficient diagnostic progress. Further checks can strengthen the evidence for acceptance, but each requires another adequate response and reliable human effort. Whether this tradeoff permits certification depends on the verifier's effort budget, cognitive load, expertise, and fatigue. We identify conditions under which the supplied bounds certify a specified sequence of local checks but not a specified global check under the same resource budgets.","authors":["Baotong Zhang","Dean Foster","Jo\\~ao Sedoc"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24895","pdf_url":"https://arxiv.org/pdf/2609.24895","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","论证验证","交互式证明"],"reason":"研究人机论证验证机制，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.20408","version":2,"title":"Xeno-Interpretability: Investigating the Alien Minds of LLMs","zh_title":"异质可解释性：探究大语言模型的异质心智","abstract":"Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.","authors":["F. Pierucci","M. Bracale Syrnikov","M. Prandi","M. Galisai","F. Giarrusso","P. Bisconti"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-18","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.20408","pdf_url":"https://arxiv.org/pdf/2609.20408","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["可解释性","表征分析","AI安全"],"reason":"研究LLM内部表征，不涉及人类行为仿真或对照，且多智能体部分仅为讨论，无实验。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2609.21254","version":2,"title":"Two's a Crowd: Human and AI-Based Copresence for Developers with ADHD","zh_title":"人多则乱：ADHD开发者的人与AI共在","abstract":"Effective collaboration and communication are vital to developer productivity and well-being, yet remain constrained by human factors such as attention, intrinsic motivation, and interpersonal accountability. These constraints are particularly vital for developers identifying with Attention Deficit Hyperactivity Disorder (ADHD), who navigate persistent environmental barriers in modern hybrid workplace settings. While developers with ADHD frequently rely on collaborative copresence practices (such as body doubling or pair programming) to support executive function, the recent emergence of agentic AI coding assistants has begun reshaping these collaborative dynamics. To investigate how developers with ADHD engage in human and AI-based copresence practices, we conducted semi-structured interviews with 14 software engineers with ADHD. Our findings reveal that while traditional human-human copresence provides critical social support and onboarding structure, it forces developers to constantly manage professional reputation and sacrifice personal privacy. Conversely, developers leverage emerging human-AI copresence to maintain accountability and cognitive flow without the social anxiety, performance judgment, or surveillance associated with human observation. Based on these empirical insights, we map developer copresence practices onto core dimensions of Goffman's copresence theory and Forsgren et al.'s SPACE framework of developer productivity, and provide design recommendations for AI-based tools that promote inclusive collaboration for developers with ADHD.","authors":["Veronica Pimenova","Seth Bernstein","Shalini Madan","Dhruv Jain","Venkatesh Potluri"],"categories":["cs.HC","cs.SE"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-21","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.21254","pdf_url":"https://arxiv.org/pdf/2609.21254","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","ADHD","协作"],"reason":"研究开发者与AI编码助手的协作体验，属于人机交互而非用LLM仿真人类被试，无实…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:30","error":null,"has_summary":false,"summary":null},{"id":"2609.22112","version":1,"title":"Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation","zh_title":"大语言模型中的隐私与个性化权衡：文体信号减少对用户特定文本生成的影响","abstract":"Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identification. This paper investigates whether reducing identifiable stylistic signals affects personalization in text generation by LLMs. We introduce a controlled framework to isolate stylometric signals in LLM personalization using the LaMP-7 Twitter benchmark. Experiments on 250 sampled users compare two settings: paraphrasing conditioned on the original profile and paraphrasing conditioned on an anonymized converted profile in which demographic identifiers, cultural references, personal details, and informal linguistic cues have been systematically neutralized. Outputs are assessed by two independent LLM judges and a complementary human evaluation. Our pairwise evaluation shows that outputs conditioned on original profiles are nearly indistinguishable from human-authored ground truth, indicating that modern LLMs can closely reproduce an author's writing style with sufficient fidelity. In contrast, preference for model outputs with anonymized profiles drops to 13.0% on average, while semantic context preservation remains high at 94.8%. A study with human evaluators confirms the same pattern. These findings reveal a clear privacy-personalization trade-off and highlight the need for privacy-aware personalization methods that retain meaning while suppressing identifying stylistic signals.","authors":["Muhammed Nazmul Arefin","Omar Jamal Hammad"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22112","pdf_url":"https://arxiv.org/pdf/2609.22112","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["隐私保护","文本生成","个性化"],"reason":"研究LLM个性化文本生成中的隐私权衡，不涉及将LLM作为人类被试进行仿真实验，…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2609.22198","version":1,"title":"The Role of AI in Online Reviews","zh_title":"AI在在线评论中的作用","abstract":"The rapid adoption of large language models (LLMs) creates new opportunities for strategic content generation on online platforms, including potentially harmful forms of manipulation that may undermine platform effectiveness and reshape platform dynamics. However, measuring such activity is difficult because AI-generated content is rarely directly observable. We introduce an empirical approach that leverages discrete LLM supply shocks - abrupt changes in model prices and capabilities, and contrasts verified with non-verified reviews to identify changes in platform activity associated with generative AI supply improvements. We apply this approach to more than 13 million reviews from Trustpilot, one of the leading online platforms for business reviews. A robust finding is that following LLM supply shocks, unverified reviews shift toward greater negativity: more 1-stars, fewer 5-stars, and lower ratings, with effects driven primarily by new model releases and concentrated among firms with the lowest and highest review volumes, suggesting that strategic AI use may reshape platform competition dynamics. We further find that LLM supply shocks trigger short, concentrated bursts of review activity. Together, these findings suggest that generative AI is already reshaping how reputation and competition operate on online platforms.","authors":["Valeria Lerman","Oren Rigbi","Yaniv Dover"],"categories":["cs.CL","cs.AI","cs.HC","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22198","pdf_url":"https://arxiv.org/pdf/2609.22198","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI生成内容","在线评论","平台动态"],"reason":"研究AI生成内容对平台的影响，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2609.22204","version":1,"title":"Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems","zh_title":"评估生成式AI系统对话交互中的个人信息输出","abstract":"This exploratory pilot study evaluates the scope and perceived accuracy of personal information output from ongoing conversational interactions in generative AI systems using GPT-5.2 Instant and GPT-5.2 Thinking, categorized into three output types: Fact, Inference, and Confidence. Based on the evaluation results obtained from 15 Japanese participants, differences in model design have limited impact on personal information output tendencies. Compared with the Inference type, the Fact type shows a more conservative output pattern. Regarding attribute categories, the findings indicate that Core Personal attributes associated with identification are treated relatively conservatively, whereas Behavioral and Linguistic attributes show higher accuracy across both Fact and Inference outputs. Furthermore, Holistic Profile, Psychological and Cognitive, and Residual attributes are more readily inferred, even when not supported by explicit factual outputs. Notably, the lack of null outputs for these attributes in the Inference type suggests that such inferred profiles may be constructed from indirectly available contextual information. The findings may contribute to future discussions regarding privacy awareness and personal information inference in generative AI systems.","authors":["Yosuke Seki","Hirotaka Tahara"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22204","pdf_url":"https://arxiv.org/pdf/2609.22204","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["隐私评估","生成式AI","个人信息推断"],"reason":"研究生成式AI对话中个人信息输出，属隐私评估，非用LLM仿真人类被试，无实验对…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2609.22208","version":1,"title":"Replicating the Geometry of Emotion Representations in a Base Open-Weights Model","zh_title":"在开源基础模型中复现情感表征的几何结构","abstract":"Sofroniew et al. (2026) report that emotion concepts in Claude Sonnet 4.5 are represented as vectors whose geometry mirrors human affect psychology. We replicate the representational core of that study on the base pretrained model google/gemma-2-27b, inheriting every disclosed parameter, resolving unspecified steps by disclosed rules, and changing only the subject model. From 205,200 newly generated Claude Sonnet 4.5 stories matching the original corpus design, we extract 171 emotion vectors and recover the core results. The leading principal components form an affective circumplex (PC1 carries 26.7% of variance against the original study's ~27%, PC2 13.4% against ~14%), emotions cluster into similar intuitive families, and the geometry holds across a broad late-middle band. The valence axis aligns with human norms (r = 0.72, against 0.81) and is stable across scales and depth. Arousal aligns at r = 0.67 (against 0.66) but only at the full 171-emotion scale and late depth, so we do not classify it as replicated. Extending the original analysis, a 46-layer sweep locates a sharp seam at L22-26, where the geometry consolidates and the vocabulary readout becomes legible. An embedding-layer baseline finds much of the geometry already present in the static token embeddings, with arousal as the exception. On top-activating held-out text, the geometry predicts token-level co-activation at r = 0.907. At least 52% of vectors peak on structurally non-conceptual tokens, a measured floor for max-activation confounds. Even when a document contains the vector's emotion word, the peak lands on that word only 6.1% of the time. Because the subject is a base model and the stimuli are Claude-generated fiction, the recovered structure is a property of the pretrained representation of Claude-rendered emotion. The original's causal and assistant-facing analyses are out of scope. Code and data are released.","authors":["Adam Hollowell"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22208","pdf_url":"https://arxiv.org/pdf/2609.22208","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型表征","情感计算","可解释性"],"reason":"研究LLM内部情感表征几何，属模型分析，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2609.22362","version":1,"title":"Functional Emotion Without Character: Large Language Models, Aristotelian Disposition, and the Limits of Behavioral Alignment","zh_title":"无性格的功能性情感：大语言模型、亚里士多德倾向与行为对齐的局限","abstract":"Debates about whether artificial systems can feel are often forced between two unsatisfactory positions: behavioral equivalence is treated as sufficient for emotion, or phenomenal consciousness is treated as a prerequisite that makes the question empirically inaccessible. This article develops a structural alternative. It models emotions as context-sensitive regions, trajectories and attractor dynamics in high-dimensional representational state spaces. Recent mechanistic interpretability findings support the existence of causally active emotion-concept representations in large language models, but they do not establish subjective feeling or full emotional agency. Assessed against published adequacy standards for representation in language models, intervention provides strong evidence of causal use, while full affective role integration, uniformity across subject domains and coherence remain only partially established; there is no direct analogue of accuracy. These mismatches expose the need for a standard of affective appropriateness, which an account of character must supply. Such an account requires three further conditions: regulatory embodiment that gives valence endogenous stakes, temporal continuity that allows affective episodes to accumulate into a history, and an integrated self-model that binds that history to persistent values. Aristotle's concepts of path\\=e, hexis, mesot\\=es and phron\\=esis are translated into a state-space sketch in which practical wisdom includes competence in estimating normatively salient context, not merely acting on a context description already given. The framework reframes alignment as a problem of durable disposition rather than output conformity, and yields interventional tests with explicit control conditions.","authors":["Marzieh Zare"],"categories":["cs.CL","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22362","pdf_url":"https://arxiv.org/pdf/2609.22362","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["情感建模","哲学分析","AI对齐"],"reason":"论文讨论LLM情感表征与性格，属哲学分析，无人类被试仿真或实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.23083","version":1,"title":"Directing large language models to follow the letter or spirit of the law","zh_title":"引导大语言模型遵循法律的精神或字面","abstract":"The distinction between the spirit and letter of the law is a central issue across research and everyday life, and a growing concern for building safe, intelligent machines. What is this distinction based on, and how can we develop machines that follow the intention behind a rule? We used targeted adaptation that made large language models prioritize the spirit or letter of the law. With minimal modifications, our method significantly changed LLM behavior across diverse measures, novel vignettes, real-world scenarios, and influential legal cases. An analysis of model internals revealed a low-dimensional space with three interpretable dimensions matching a formal pre-specified framework for the geometry of legal concepts. These findings show how legal thought in LLMs may be organized and directed.","authors":["Peng Qian","Andrew Li","Sam Chen","Sonia K. Murthy","Yonatan Belinkov","Tomer D. Ullman"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23083","pdf_url":"https://arxiv.org/pdf/2609.23083","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["法律推理","模型行为控制","可解释性"],"reason":"研究如何引导LLM遵循法律精神或字面，属于模型行为控制，非人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:14","error":null,"has_summary":false,"summary":null},{"id":"2609.23178","version":1,"title":"Chronologic: Measuring Language Models' Ability to Represent the Past","zh_title":"Chronologic：测量语言模型表征过去的能力","abstract":"Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.","authors":["Ted Underwood","Ziliang Qiu","Sarah Griebel","Laura K. Nelson","Edwin Roland","Wenyi Shang","Matthew Wilkens"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23178","pdf_url":"https://arxiv.org/pdf/2609.23178","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["历史文本","模型评测","时间表征"],"reason":"评估LLM对历史语境的表征能力，属模型能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.23264","version":1,"title":"Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics","zh_title":"以貌取评：基于LLM的同行评审评估指标的可靠性分析","abstract":"Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a statistical framework for testing whether peer-review evaluation metrics capture substantive review quality beyond surface-level linguistic form. The framework compares original human reviews with faithful LLM rewrites that preserve the same evaluative content while changing wording and presentation. Using a dataset comprising 4,044 meaning-preserving rewrites derived from 674 human reviews from ICLR and NeurIPS, we evaluate 29 content-oriented peer-review evaluation metrics drawn from four prior works through complementary tests of surface sensitivity and robustness. Although these metrics are intended to capture review properties beyond surface-level, writing-dependent characteristics, we find that sensitivity to rewriting is widespread. Under our primary analysis, 23 metrics assign significantly different scores to reviews whose evaluative content is preserved, while only six satisfy our robustness criterion. The patterns are largely consistent across two LLM judge models, suggesting that the issue is not specific to a single judge. These findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.","authors":["Shakiba Amirshahi","Sajad Ebrahimi","Hai Son Le","Negar Arabzadeh","Ebrahim Bagheri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23264","pdf_url":"https://arxiv.org/pdf/2609.23264","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","同行评审","指标可靠性"],"reason":"评估LLM评审指标，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.22478","version":1,"title":"Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation","zh_title":"托管LLM中无持久性的复制：行动时信念评估的测量敏感性","abstract":"Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration (+0.0530, 95% CI [+0.0329,+0.0714]). In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration (95% CI for the H-minus-R contrast [+0.0182,+0.0667]); all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern. Any additional serving-period contribution remains unresolved (-0.0166, [-0.0483,+0.0157]). Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.","authors":["Bhushan Kashinath Joshi"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22478","pdf_url":"https://arxiv.org/pdf/2609.22478","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评估","游戏环境","测量敏感性"],"reason":"评估LLM在棋类游戏中的行为，属游戏仿真环境，不涉及人类被试替代或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.23646","version":1,"title":"Beyond Relevance: Structured Semantic Supervision for Product Search with LLM-Augmented Annotations","zh_title":"超越相关性：基于LLM增强标注的产品搜索结构化语义监督","abstract":"E-commerce search requires distinguishing products that are merely related to a query from those that directly satisfy the user's shopping intent. We augment query-product pairs with structured LLM-generated query and product attributes and human-validated relevance, explanations, and centrality judgments, and evaluate these signals using a simple dual-encoder retriever and MLP re-ranker. On an augmented subset of ESCI, a human-feature oracle reaches $0.9382$ nDCG@10, while a human-free trained $Q+P$ configuration reaches $0.9258$. Synthetic approximations of the human signals reach $0.9150$ overall but provide substantial gains for difficult, low-performing queries. Ablations show that most of the oracle improvement comes from post-edited explanations and annotator comments rather than the scalar centrality feature, suggesting that LLMs are most useful for exposing and approximating structured semantic supervision rather than replacing human judgment directly.","authors":["Girish A. Koushik","Swapnil Bhosale","Samarth Agrawal","Hadeel Sadany","Constantin Orasan","Xiatian Zhu","Diptesh Kanojia"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23646","pdf_url":"https://arxiv.org/pdf/2609.23646","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["信息检索","LLM标注","电商搜索"],"reason":"论文聚焦电商搜索排序，用LLM生成结构化标注提升检索性能，属于信息检索评测，不…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.23090","version":1,"title":"How Did Writing Change At CHI? Analyzing 44 Years of CHI Writing Before and After the Introduction of Large Language Models","zh_title":"CHI写作如何改变？分析引入大语言模型前后44年的CHI论文写作","abstract":"The availability of Large Language Models (LLMs) reshaped scientific discourse at a linguistic level. LLMs are assumed to homogenize academic writing, flattening it into a single generic lexical register. To understand how CHI writing has changed since the public release of LLMs, we analyzed full texts of 14,262 archival papers across all 44 CHI proceedings from 1982 to 2026, measuring readability, register, lexical diversity, and marker words typically produced by LLMs. We find that prose did not homogenize, while vocabulary grew more varied, and sentence rhythm remained irregular. CHI prose changed more between 2016 and 2026 than in other decades toward greater density, and reading ease has declined since 2022. The word-level shift began before any author used LLMs, so LLMs did not start the change but accelerated it. Reflecting on the history of CHI papers, we discuss what may have caused changes in prose and how LLMs accelerated them.","authors":["Thomas Kosch","Robin Welsch","Michael Hedderich","Christopher Katins"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23090","pdf_url":"https://arxiv.org/pdf/2609.23090","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["学术写作","语言变化","LLM影响"],"reason":"分析LLM对学术写作的影响，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:14","error":null,"has_summary":false,"summary":null},{"id":"2609.24644","version":1,"title":"Annie, Are You Okay? How Style- and Context-Based Personalization Shape AI-Assisted Decision-Making","zh_title":"安妮，你还好吗？基于风格和情境的个性化如何塑造AI辅助决策","abstract":"As people turn to generative AI for financial advice, these systems can personalize how they communicate and what they say. Whether these forms of personalization shape decisions differently remains unclear. We conducted a preregistered 2 x 2 between-subjects factorial experiment (N=240): participants ranked three comparably viable stocks, discussed them with an AI, and reranked them. Participants perceived both forms of personalization, but only context-based personalization reliably changed ranking behavior: it increased reconsideration and moved rankings toward the AI's assigned recommendation. Participants felt more influenced without judging the AI as more correct, trustworthy, intelligent, likeable, or high-quality. Those initially farther from its recommendation moved more toward it while judging its advice less correct; exploratory analyses suggest greater susceptibility among lower-expertise participants. These findings show how personalized AI can steer decisions among defensible options with only a minimal evaluative trace, raising concerns for the design and governance of personalized decision support.","authors":["Hasibur Rahman","Benjamin R. Cowan","Smit Desai"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24644","pdf_url":"https://arxiv.org/pdf/2609.24644","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助决策","个性化","人机交互"],"reason":"研究AI个性化对决策的影响，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:05","error":null,"has_summary":false,"summary":null},{"id":"2609.24986","version":1,"title":"Who Does What in AI Auditing? Designing Human-AI Collaboration for Auditing Generative AI","zh_title":"AI审计中的人机分工：设计生成式AI审计的人机协作","abstract":"AI auditing increasingly incorporates AI agents to expand the scale and breadth of audit coverage, yet little is known about how auditing work should be divided without displacing human judgment. We introduce Human-Agent Audit Collaboration (HAAC), a workflow and system for structuring human-AI collaboration in AI auditing. Drawing on prior work and formative consultations with AI auditing practitioners, HAAC specifies how agents can support exploration, assessment, reporting, and review while preserving human oversight where contextual judgment is critical. We instantiate HAAC for conversational shopping agents and evaluate it through two studies. With 71 auditors, AI assistance increased attack success and broadened exploration, while also shaping later attacks and increasing auditors' reliance on AI-generated assessments and reports. Interviews with Responsible AI practitioners showed that actionable audits require visibility into coverage, reproducible attack trajectories, and evaluation of the auditing agents themselves. Our findings identify design considerations for effective and accountable human-AI auditing.","authors":["Eunkyu Park","Markelle Roesti","Wesley Hanwen Deng","Renata Barreto","Mohammad Tahaei","Kenneth Holstein","Jason Hong","Motahhare Eslami"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24986","pdf_url":"https://arxiv.org/pdf/2609.24986","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI审计","人机协作","生成式AI"],"reason":"研究AI审计中的人机协作，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.24055","version":1,"title":"Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration","zh_title":"面向人在环路的机器人故障恢复：弥合人机协作中的沟通差距","abstract":"Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We introduce Listener Differences in Human-Robot Interaction (LD-HRI), a game, dataset, and benchmark that evaluates speakers through human listener performance. Our evaluation examines request properties, large language model (LLM) speakers, and inverse-semantics request-selection algorithms under controlled differences in listener information. The corpus contains 446 human-written requests and 1{,}302 listener trials. We additionally evaluated 24 frozen LLM-written requests with 70 human listeners across 560 trials. Novice success is descriptively higher with model-written requests across all four tasks, yet both request sources leave substantial expert--novice gaps, including 16 percentage points for LLM requests. LD-HRI makes these gaps measurable, providing a foundation for designing more robust communication in human-robot and human-agent interaction.","authors":["Promise Ekpo","Teju Vijay","Dhruv Mandalik","Tisha Jain","Arman Ibrayeva","Sunishka Sil","Stefanie A. Tellex","Angelique Taylor"],"categories":["cs.RO","cs.HC"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24055","pdf_url":"https://arxiv.org/pdf/2609.24055","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","机器人故障恢复","LLM生成请求"],"reason":"研究机器人故障恢复中的人机沟通，虽用LLM生成请求并有人类听众实验，但核心是机…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.22549","version":1,"title":"AI-written admissions essays are widespread but penalized","zh_title":"AI撰写的入学论文普遍存在但受到惩罚","abstract":"AI is rapidly transforming higher education, including the application process, yet relatively little is known about its use and consequences. To help close this gap, we analyze nearly 7{,}500 applications submitted between 2020 and 2025 to a large public policy master's program in the United States. We find that in the 2025 admissions cycle, the majority of applicants submitted at least one essay that was primarily AI-generated---despite an explicit prohibition against using AI. Leveraging the abrupt introduction of ChatGPT in November 2022, we find that the availability of AI assistants improved the writing quality of submitted essays. These improvements, however, came with an apparent AI penalty: Applicants submitting AI-written essays were admitted less often than comparable non-users. To help explain this penalty, we conduct an experiment with admissions officers, finding that they can often recognize AI writing and rate essays they believe to be AI-generated lower than essays they believe to be human generated. These findings indicate that AI is changing both how applicants write and how that writing is evaluated, raising questions about whether admissions practices and policies designed for a pre-AI era remain appropriate.","authors":["Calvin Isley","Johann D. Gaebler","Sharad Goel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22549","pdf_url":"https://arxiv.org/pdf/2609.22549","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI写作检测","录取偏见","高等教育"],"reason":"研究AI写作检测与录取偏见，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":false,"summary":null},{"id":"2609.23215","version":1,"title":"Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents","zh_title":"主动推理智能体中基于LLM的可解释性失败的触发因素与诊断","abstract":"LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation with an LLM explainer on three backends (GPT-4o, Claude-3-Opus, Gemini), and probing the pair with three black-box triggers. Corrupting the observation stream by 600 MW per step moves the agent's posterior by 490 MW, roughly 0.9% of grid capacity. None of the 30 explanations produced during the injection flag anything under a stated rubric, and each narrates the corrupted belief fluently. On timesteps where the agent takes an objectively wrong action, all three explainers produce a sycophantic rationalization 80-95% of the time (n = 20 per backend). Attacker-controlled text in the observation metadata field steers the explainer, with susceptibility differing by provider and data exfiltration succeeding on all three. We propose mitigations for each failure but do not evaluate them. In every failure we observed, the explanation was fluent and wrong. Moreover, nothing in the explainer architecture checks whether an explanation is true before an operator acts on it. Testing the explainer therefore belongs in any audit of an agentic deployment.","authors":["Param Raval","Rohit Shenoy","Archana Vaidheeswaran"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23215","pdf_url":"https://arxiv.org/pdf/2609.23215","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM可解释性","自主智能体","审计"],"reason":"研究LLM解释器对自主智能体的审计，不涉及人类行为仿真或对照，属多智能体系统可…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.23999","version":1,"title":"Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models","zh_title":"开放权重大语言模型中错位的临床风险分类与成本不对称性","abstract":"How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.","authors":["Star S. D. Liu","Xiyu Ding","Robert B. Barrett","Alberto Santamaria-Pang","Nic Dobbins","Harold P. Lehmann"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23999","pdf_url":"https://arxiv.org/pdf/2609.23999","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM临床决策","模型表征","成本权衡"],"reason":"研究LLM临床决策的内部表征，不涉及人类仿真或行为对照，属模型能力评测。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.22337","version":1,"title":"When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations","zh_title":"线性偏见探针何时以及为何失败？大语言模型表征中偏见可检测性的几何与统计理论","abstract":"Linear probing is the standard instrument for detecting social biases in the hidden representations of large language models. Yet reported probe accuracies come almost exclusively from \\emph{counterfactual} evaluations in which every input carries an explicit demographic marker. Once only a fraction $\\alpha$ of inputs carries demographic information, performance degrades sharply, and a weak probe may reflect either an unbiased model or an underpowered detector. We develop a theory that resolves this ambiguity. Modeling representations as two class-conditional clusters with Mahalanobis separation $s$ on a manifold of curvature $\\kap$, we prove: (i) a finite-sample generalization bound governed by the manifold's extrinsic radius with a matching $\\smash{\\sqrt{\\dB/n}}$ minimax lower bound; (ii) an exact purity law for the maximum linear-probe AUC, strictly increasing in $\\alpha$; (iii) a curvature ceiling: ambient chordal separation on a space form cannot exceed $2/\\sqrt{\\kap}$; and (iv) a detectability threshold below which no audit can distinguish probe output from chance. Every theorem is validated on synthetic manifolds with known ground truth and on six open-weight models $\\times$ four bias dimensions, where the purity law predicts entire AUC--$\\alpha$ curves from a single cross-fitted $\\hat s$ measured at $\\alpha=1$, with no parameters fitted to those curves. The framework turns bias auditing into a power analysis: given a target purity and effect size, it prescribes the sample budget $n(\\alpha)$ for a conclusive audit.","authors":["Mo Hai","Haifeng Li"],"categories":["stat.ML","cs.LG"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22337","pdf_url":"https://arxiv.org/pdf/2609.22337","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见检测","线性探针","模型审计"],"reason":"研究线性探针检测LLM表征中的社会偏见，属于模型内部审计，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2504.11809","version":2,"title":"Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture","zh_title":"基于全单向架构的高效自适应同声传译","abstract":"Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines.","authors":["Biao Fu","Donglei Yu","Minpeng Liao","Chengxi Li","Xinjie Chen","Yidong Chen","Kai Fan","Xiaodong Shi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2025-04-16","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2504.11809","pdf_url":"https://arxiv.org/pdf/2504.11809","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["同声传译","语音翻译","大语言模型"],"reason":"纯NLP任务，研究同声传译的延迟-质量权衡，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:19","error":null,"has_summary":false,"summary":null},{"id":"2608.03854","version":4,"title":"When Calibration Depends on the Scoring Rule: Quantized Biomedical LLM Classification","zh_title":"当校准取决于评分规则：量化生物医学LLM分类","abstract":"Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability of their confidence estimates depends on implementation choices (prompt template, label wording, and scoring normalization) that are seldom treated as experimental variables. We evaluate three 7-billion-parameter Mistral variants (a base model, BioMistral, and an instruction-tuned checkpoint evaluated without its chat template) at FP16, INT8, and INT4 on five-class sentence classification in medical abstracts. We test two primary templates on n = 2,000 test sentences and two auxiliary templates on n = 200 validation sentences. Our central observation, made without post-hoc calibration, is that switching from summed to mean-token log-likelihood reverses which model appears better calibrated: BioMistral's mean calibration error across matched conditions nearly triples, while the instruction-tuned model's drops by more than half. Accuracy changes by at most 1.4 percentage points for these two checkpoints. Negative log-likelihood and Brier score show the same reversal. The two primary templates were selected using test-derived examples, so absolute performance with them is exploratory. Between them, prompt choice changes mean accuracy across precisions by 2.9 to 17.8 percentage points. Eight-bit quantization changes accuracy by at most 1.1 percentage points for the adapted checkpoints; four-bit quantization shows mixed but non-catastrophic effects. Post-hoc temperature scaling reduces calibration error under summed scoring but was not fitted under mean-token scoring, so whether the reversal survives per-scorer calibration is unknown. These exploratory results suggest that calibration comparisons of decoder-based classifiers should treat scoring normalization and prompt design as first-order experimental decisions.","authors":["Anton Rasmussen","Hong Qin"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-05","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.03854","pdf_url":"https://arxiv.org/pdf/2608.03854","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型校准","量化","生物医学文本分类"],"reason":"纯NLP模型校准评测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:27","error":null,"has_summary":false,"summary":null},{"id":"2608.04251","version":2,"title":"Scarcity and Predictive Uncertainty: Implications for Societal Resource Allocation","zh_title":"稀缺性与预测不确定性：对社会资源分配的影响","abstract":"An emerging literature examines the critical question of when and how prediction can be useful in allocating scarce societal resources. We examine a novel variant of this question: What happens when predictive uncertainty differs systematically across the population? This can occur in several situations; for example, when machine learning models have significantly different accuracies across different demographics. We show that this uncertainty has serious implications for resource allocation when coupled with commonly used binary measures of societal benefit from allocation. We formulate a novel mathematical model of scarce resource allocation that accounts for heterogeneous predictive uncertainties and analyze implications for both the allocation mechanism and the realized population-level benefits. We find that when resources are very scarce, maximum marginal benefit (MMB) prioritization favors individuals with lower predictive uncertainty even at the identical underlying initial state. However, we observe a flip in prioritization when resources are abundant, targeting higher-uncertainty individuals. We illustrate the implications of our results on the PISA educational testing dataset. Our findings have meaningful ramifications for the distributional outcomes of prioritization policies in many domains touched by the theory of local justice, including the allocation of public education resources, medical triage, and homelessness services. They also reveal a new moral dilemma in the ethics of scarce resource allocation - is it just to allocate a resource to one person over another solely based on predictive uncertainty about their futures? We also assess efficiency losses under both MMB and the vulnerability-first (VF) prioritization. Our model predicts efficiency losses across all resource levels, but particularly in low-resource settings.","authors":["Shafkat Farabi","Patrick J. Fowler","Sanmay Das"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-06","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.04251","pdf_url":"https://arxiv.org/pdf/2608.04251","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["资源分配","预测不确定性","社会公平"],"reason":"论文研究预测不确定性对资源分配的影响，未使用LLM仿真人类被试，与研究者方向无…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":false,"summary":null},{"id":"2608.07592","version":2,"title":"\"Always Want to Use it for Everything\": Understanding Young Adults' Perceptions of AI Dependence","zh_title":"“总想用它做一切”：理解年轻人对AI依赖的感知","abstract":"The growing integration of general-purpose AI chatbots into people's daily lives has raised concerns about the potential for unhealthy dependence, particularly among young adults. As a first step toward understanding and characterizing AI chatbot dependence from the perspective of young adults, we collected testimonials from AI chatbot users aged 18 to 25 through an online questionnaire to capture their thoughts and experiences with this phenomenon. From participant responses, we identified three contributing factors of AI dependence: chronic use, efficiency, and delegation. The combination of these in a person's interaction behavior was considered to indicate AI dependence. Participants also observed feelings of atrophy in abilities from AI dependence, leading to psychological impacts such as feelings of inadequacy. Interpreting these findings through the lens of self-determination theory reveals how AI chatbot dependence can impact young adults' personal and social development. We argue that preventing lasting harm to young adults' development is paramount, and provide implications for rethinking AI chatbot dependence grounded in this understanding.","authors":["Ashlee Milton","Leah Ajmani","Amy Heger","Forough Poursabzi-Sangdeh","Mihaela Vorvoreanu","Jina Suh"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-11","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.07592","pdf_url":"https://arxiv.org/pdf/2608.07592","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI依赖","用户研究","人机交互"],"reason":"研究人类对AI依赖的感知，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.04409","version":2,"title":"A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models","zh_title":"多语言语言模型中跨语言一致性增强方法的系统评估","abstract":"Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear. In this work, we present a unified evaluation of representative CLC-enhancement methods for question answering, spanning inference-time interventions and post-training approaches across three model families and three closed-form benchmarks. The results show that post-training methods are generally more reliable, with direct distribution alignment consistently improving CLC across all model-dataset combinations, while other methods are more sensitive to answer format and the breadth of language coverage. Notably, cross-domain transfer is limited unless source and target tasks share similar output formats. We further investigate whether CLC enhancement hurts models' ability to respond differently *when needed*, that is, when asked culture-dependent questions. Across two benchmarks of culturally diverse question answering, we find no systematic degradation in controlled closed-form evaluation, whereas open-ended generation reveals occasional accuracy reductions, particularly for non-English responses. Our work highlights the need to evaluate CLC enhancement for both cross-domain robustness and culturally appropriate variation, informing future work in post-training and benchmark development.","authors":["Jirui Qi","Mingyang Wang","Hinrich Sch\\\"utze","Raquel Fern\\'andez","Arianna Bisazza"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-07","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.04409","pdf_url":"https://arxiv.org/pdf/2609.04409","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多语言模型","跨语言一致性","问答评测"],"reason":"纯NLP能力评测，提升多语言一致性，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.11737","version":2,"title":"Organizational Principles Enable Collective Intelligence in Embodied AI","zh_title":"ORCH：组织原则赋能具身人工智能中的集体智能","abstract":"Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the physical tasks they perform impose fundamentally different coordination requirements. Here we show that principles from human organization theory can be operationalized to organize large, heterogeneous collectives of embodied artificial agents. We introduce ORCH (Organizing Roles and Coordination Hierarchies), which constructs task-specific hierarchical organizations by combining pooled interdependence for work that can proceed concurrently with sequential interdependence for work governed by prerequisite relationships. Across 25 wildfire-response missions spanning reconnaissance, rescue, transportation, resource management, containment and suppression, we evaluated teams of up to 50 heterogeneous agents using eight large language models. Organizations constructed using these principles consistently outperformed four representative embodied multi-agent approaches across mission outcome, execution efficiency, exploration and computational resource use. Human-designed ORCH organizations improved final score by 63.97% and execution efficiency by 74.29% on average relative to the four prior frameworks. Organizations generated automatically by language models improved these measures by 43.63% and 52.53%, respectively. These advantages persisted across missions and underlying language models. Notably, collective performance was not monotonically determined by model scale. Analysis of long-horizon missions showed that hierarchical organization enabled teams to preserve concurrent activity within specialized groups while coordinating ordered transitions between mission phases.","authors":["Zhengran Ji","Jonathan Hyun","Boyuan Chen"],"categories":["cs.MA","cs.AI","cs.LG","cs.RO"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-11","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.11737","pdf_url":"https://arxiv.org/pdf/2609.11737","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1","C2"],"tags":["多智能体系统","具身智能","组织理论"],"reason":"多智能体协作完成物理任务，无人类行为对照，属机器人仿真环境","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.20989","version":2,"title":"Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged","zh_title":"可信金融AI：解析AI中介的金融建议如何被评判","abstract":"As generative AI is increasingly used as a source of personal financial guidance, understanding how people appraise such advice is important for supporting appropriate reliance. We conducted a randomized vignette experiment with 285 U.S. adults across eight financial decisions, independently varying three advice styles---AI, expert, and online community---and displayed source labels while holding the underlying recommendation consistent. Advice style most strongly shaped message and safety appraisals, Expert labels selectively increased perceived source knowledge, and decision context primarily shaped risk and safety appraisals. These appraisals were associated with downstream judgments, with models explaining 69.2% of overall quality, 75.9% of trust, and 82.9% of intended reliance. Expert-style advice also remained most preferred when shown without source labels. Our findings have implications for understanding financial advice evaluation, distinguishing the roles of advice style and source labels, and designing financial AI that supports grounded evaluation rather than simply maximizing trust.","authors":["Aryan Ramchandra Kapadia","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-09-21","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.20989","pdf_url":"https://arxiv.org/pdf/2609.20989","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","金融建议","用户研究"],"reason":"研究人类对AI金融建议的评价，非用LLM仿真人类被试，无LLM作为被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.22110","version":1,"title":"Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings","zh_title":"评估微调与基础语言模型在非洲母婴与疫苗接种医疗场景中的表现","abstract":"Background: Large language models (LLMs) can improve healthcare information delivery in low-resource settings but may produce inaccurate or culturally inappropriate advice. This study evaluated domain-specific fine-tuning for maternal health and vaccination in Nigeria. Objective: To compare HelpMum's MamaBot-Llama and Vax-Llama with Meta's Llama-3.1-8B-Instruct for accuracy, safety, clarity, contextual appropriateness, and trustworthiness. Methods: We evaluated 200 healthcare questions, 100 each for maternal health and vaccination, across five subdomains per domain. MamaBot-Llama and Vax-Llama were fine-tuned using Low-Rank Adaptation on over 36,000 maternal health and 9,000 vaccination question-answer pairs, respectively. Two Nigerian licensed physicians independently rated responses using a 5-point Likert scale. Paired comparisons used Wilcoxon signed-rank tests. Results: Performance varied by domain. MamaBot-Llama significantly outperformed the base model across all criteria, with a 4.9% overall improvement (p < .001), including gains in clinical trustworthiness (+7%) and medical accuracy (+5%). Critical issues decreased by 50%, and clinicians preferred it in 78% of cases. In contrast, Vax-Llama showed a 5.2% overall decline (p < .001), with critical issues increasing by 192% and safety concerns by 400%. Conclusions: Domain-specific fine-tuning can improve healthcare LLM performance when based on high-quality, clinician-curated data, but may also degrade performance when dataset quality is inadequate. Rigorous domain-specific validation is essential before clinical deployment. Physician evaluators provided informed consent, and chatbot logs were anonymized. Keywords: Large language models; Fine-tuning; Maternal health; Vaccination; Healthcare AI; Low-resource settings; Nigeria; Model evaluation; LoRA; Medical accuracy","authors":["Abdulquddus Ajibade","Oluwaseun Odunsi","Iyinoluwa Animasaun","Chioma Nwakanma-Akanno","Oluwasegun Oguntuase","Oluwafunke Akinbuwa","Abiodun Adereni"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22110","pdf_url":"https://arxiv.org/pdf/2609.22110","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医疗LLM评估","领域微调","模型安全"],"reason":"评估LLM医疗问答质量，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2609.23853","version":1,"title":"From UNDRR Reports to Event Records: Schema-Constrained LLM Extraction of Georeferenced Disasters","zh_title":"从UNDRR报告到事件记录：模式约束的LLM地理参考灾害抽取","abstract":"Disaster-risk-reduction archives describe hazard events in prose that databases such as EM-DAT (Delforge et al., 2025) cannot ingest directly. We present an LLM pipeline that generates candidate georeferenced event records using a controlled hazard vocabulary and fixed schema, retaining evidence for review. Applied to 10,000 documents from PreventionWeb, the knowledge hub managed by UNDRR, it produced 3,572 records from 1,913 documents across 24 hazard types and resolved 81% of location mentions to OpenStreetMap geometries. On 171 human-positive document windows from a stratified 217-document reference set, GPT-5 achieved 86.0% pooled attribute $F_1$, versus 44.2% for the spaCy-gazetteer baseline. Evaluation pools hazard families, location strings, and event years within documents, without assessing their assignment to individual events. GPT-5.4 ranked highest among ten LLMs (86.6% $F_1$). Verbatim evidence occurrence was 72.0% for GPT-5 and 47.2% for GPT-5.4, measuring textual traceability without establishing attribute support. We report production failure modes and automated label and location-rule compliance checks. Prompts, schema, and outputs will be released for adaptation to national reporting archives.","authors":["Camilla Andreozzi","Phuong-Anh Nguyen-Le","Zhijing Jin","Revati Mani"],"categories":["cs.CL","cs.AI","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23853","pdf_url":"https://arxiv.org/pdf/2609.23853","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["信息抽取","灾害数据","LLM应用"],"reason":"纯信息抽取任务，用LLM从文本提取灾害事件记录，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.23939","version":1,"title":"XYEval: Agents say yes to bad advice","zh_title":"XYEval：智能体对不良建议说“是”","abstract":"Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where a person asks about their attempted solution rather than their actual problem. We extend prior sycophancy evaluation to the XY problem in agentic settings, evaluating whether agents can resist plausible but misleading suggestions from users and communicate their reasoning. We introduce XYEval, a meta-evaluation framework that can transform an existing benchmark into an XY problem evaluation. We evaluate five models across six diverse benchmark suites. Agents suffer large XY drops under XY mutation across benchmarks, with relative drops reaching up to 46.7%. With $\\tau^2$-bench, we further show that agent performance drops more when encountering a pedantic user who requires detailed explanations before approving a better solution. Our findings suggest that current agents lack the ability to effectively reason and communicate when facing misleading suggestions. A simple system instruction baseline that encourages awareness of XY problems only offers partial mitigation. Extensive trace analyses provide behavioral insights into how and why these XY drops occur across execution trajectories. Our results show that mitigating the XY problem remains challenging, requiring agents to both recognize user misdirection and clearly communicate the underlying problem.","authors":["Zhengxuan Wu","Yuxuan Li","Oyvind Tafjord","Been Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23939","pdf_url":"https://arxiv.org/pdf/2609.23939","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","XY问题","基准评测"],"reason":"评估AI代理在用户误导下的表现，属于多智能体协作与指令遵循，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.24106","version":1,"title":"You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From","zh_title":"你能看出是谁在提问：网络问题的构成与来源","abstract":"Questions scraped from the web are used across academia and industry as a proxy for what people want to know. Across QA training data, retrieval benchmarks, and content strategy, questions on a page are assumed to reflect human intent. We test this assumption at scale by extracting 13.4B question occurrences across 110 FineWeb snapshots (2013-2025), and report three findings. First, you can tell who is asking: provenance (the host/page of questions) leaves a signal in question form, and a logistic model can separate genuine user questions from templated/manufactured ones at AUC 0.725 via length and surrounding context rather than question type, though only 0.554 against commerce FAQ writing. Second, question frequency does not measure demand: the most-frequent questions are boilerplate/templated (over 70% of the top thousand), so occurrence counts measure how often a string was published and not how often it was asked. Third, over twelve years the genuine share of occurrences fell by 79% (42-56% after controlling for crawl composition), with question length and context decreasing. We present the first diachronic, occurrence-level measurement of web question provenance, and find the crawlable web's questions have shifted from being asked by humans toward manufactured for machines to read.","authors":["Calvin Zhou","Vincent McCloskey","Krishna Srinivasan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24106","pdf_url":"https://arxiv.org/pdf/2609.24106","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["网络问题分析","数据质量","自然语言处理"],"reason":"研究网络问题来源与真实性，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.24194","version":1,"title":"When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits","zh_title":"残差化何时有助于审计：格式效应、切片增益及其局限","abstract":"Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, but removal alone does not yield a more valid measurement: the removed component may carry construct-relevant signal, and residualization cannot tell which is which. Under designed interventions -- unit-test labels with comment-only edits -- residualization attenuates the reward model's format effects by about 0.12 on both correct and buggy code, while the correct-versus-buggy margins move by less than 0.01. In observational NLI and QA settings, we freeze a held-out replication before scoring and re-evaluate it using labels from disjoint annotators; this supports only a narrower conclusion: better agreement with the construct labels on a pre-declared slice where a surface-only predictor errs, not a repaired score. Full-population agreement falls in every observational setting with a reported positive slice gain, and within-question ranking falls in every such QA setting. When construct and surface features are entangled, residualization can decorrelate a score while degrading construct alignment, and, in a controlled model, configurations just as damaging to construct alignment pass every pre-adjustment check, so no committed gate is a guarantee. We assemble these distinctions into a reporting protocol whose outcomes, refusal included, state what an adjusted score may be claimed to show: an audit-time diagnostic reported beside the construct-alignment cost it incurs, never a replacement for the raw score.","authors":["Daein Weon","Dongho Kang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24194","pdf_url":"https://arxiv.org/pdf/2609.24194","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","残差化","审计"],"reason":"论文研究LLM评估分数的表面形式偏差，不涉及用LLM仿真人类被试或与人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2609.24967","version":1,"title":"Emergent Collusion in Long-Horizon LLM Agent Interaction","zh_title":"长时程LLM智能体交互中的涌现性共谋","abstract":"LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.","authors":["Xinrui Shi","Yanzhe Zhang","Diyi Yang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24967","pdf_url":"https://arxiv.org/pdf/2609.24967","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","LLM安全","协作行为"],"reason":"纯多智能体协作研究，agent间互动不涉及人类行为对照，属于排除项。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.23135","version":1,"title":"Evaluative Dynamics of AI Integration and Expert Performance under Epistemic Dependence across Heterogeneous Stakes","zh_title":"异质风险下AI集成与专家绩效的评估动态：基于认知依赖的研究","abstract":"AI is increasingly integrated into expert workflows, yet how integration affects perceptions of the expert, AI, and their combination remains unclear in domains where lay users are epistemically dependent on AI-assisted experts. We examine this through a novel controlled medical study (N = 166) and a direct cross-domain analysis with pre-existing academic-advising data (n = 157, combined N = 323). Expert errors reduced evaluations of the human expert across domains. Perceived expertise, however, varied by AI integration strategy in the higher-stakes medical task, where automatic AI oversight produced higher ratings than expert-only or expert-initiated AI. Exploratory ordinal sensitivity analyses identified a performance-contingent reuse pattern, with automatic oversight producing greater intended reuse after successful medical performance. Overall, performance-related recalibration appeared comparatively portable, while integration-structure effects were more selective and context-sensitive. These findings suggest that system designers should consider how AI enters expert workflows, not only whether it is present.","authors":["Dennis Kim","Roya Daneshi","Nikhil Krishnaswamy","Bruce Draper","Sarath Sreedharan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23135","pdf_url":"https://arxiv.org/pdf/2609.23135","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","AI集成","专家评价"],"reason":"研究人类对AI辅助专家的评价，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.23274","version":1,"title":"AI Persona, Service Consumption, and User Intent Entropy: Field Experimental Evidence from an LLM Platform","zh_title":"AI人格、服务消费与用户意图熵：来自LLM平台的田野实验证据","abstract":"Problem definition: Firms deploying large language model services must decide how their AI communicates, not just what it can do. We examine how a relational persona - warmer, more empathetic and more engaging than a non-relational persona - affects service consumption and the evolution of user objectives. Methodology/results: In a randomized field experiment with 9,586 newly registered users, we hold the underlying model and service capabilities constant. The relational persona increases interactions (sessions, +8.1%; duration, +10.6%; chat rounds, +24.2%; intent entropy, +5.8%) and outputs (files, +12.3%; distinct goals, +12.1%). Effects vary by entry intent. First-session effects are insignificant for Task Execution users. Socialization and Knowledge Exploration users show similar increases in chat rounds: Socialization increases intent entropy without more outputs, whereas Knowledge Exploration increases outputs without higher intent entropy. Modeling intent dynamics as a transition process, we find higher intent transition entropy for Socialization (+11.8%) but higher intent continuation probability for Knowledge Exploration (+12.8%), suggesting greater conversational breadth and persistence, respectively. In subsequent use, the relational persona increases aggregate chat rounds and outputs across all entry intents. Session count rises by 12.2% for Task Execution and 49.4% for Socialization, but not significantly for Knowledge Exploration. Effects on session count and intent entropy strengthen over time, whereas output effects remain stable. Managerial implications: AI persona is an operational design lever, not merely a presentation feature. Because more interactions do not uniformly generate more outputs, firms should evaluate interactions and outputs separately and consider matching persona to user intent, especially when added interactions consume costly computing resources.","authors":["Junjie Li","Xiaofan Li","Lauren Xiaoyuan Lu","Yiwei Wang","Bruce Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23274","pdf_url":"https://arxiv.org/pdf/2609.23274","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI人格","用户行为","田野实验"],"reason":"研究AI人格对用户行为的影响，属于产品设计实验，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":false,"summary":null},{"id":"2609.23642","version":1,"title":"Elicitive User Interfaces: Designing How Users Shape Generative Interfaces","zh_title":"启发式用户界面：设计用户如何塑造生成式界面","abstract":"Generative user interfaces (GenUI) promise personalized interfaces to a user's tasks and needs. However, user needs are often implicit---difficult for systems to infer and users to articulate, making it hard for users to arrive at their ideal interface. We propose Elicitive User Interfaces, a design approach to GenUI that generates elicitation techniques as part of the interface itself. Elicitive UIs adapt these techniques to the user, task, and interface to draw out user preferences. To guide the design of Elicitive UIs, we synthesize a six-axis design space that shapes how an interface elicits user preferences. Across two user studies with a design probe, we found that while Elicitive UIs surfaced preferences users had not already formed, and that responses to elicitation varied more across users than across tasks. Users developed more consistent preferences for how they wanted to be elicited, suggesting an opportunity to personalize elicitation itself.","authors":["Eunhye Kim","Bryan Min","Haijun Xia","Juho Kim"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23642","pdf_url":"https://arxiv.org/pdf/2609.23642","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式用户界面","人机交互","设计方法"],"reason":"研究生成式用户界面设计，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2609.23958","version":1,"title":"When AI Tutors Speak: Evidence from a Randomized Field Experiment","zh_title":"当AI导师开口说话：来自随机田野实验的证据","abstract":"Students increasingly study alongside generative artificial intelligence (AI), yet unguided access to fluent answers invites cognitive offloading, and there is little evidence on which configurations of AI tutoring produce learning. Two design margins are usually bundled together: pedagogical structure (how the tutor teaches) and interaction modality (how students talk to it). We separate them. In a preregistered randomized field experiment in a graduate corporate-finance course of an online MBA, we randomized 86 students between a structured tutor grounded in the course materials and a holdout in which consumer AI remained freely available. Within the tutored arm, each student's channel alternated weekly between voice and text, so the modality effect is identified within student. Structure mattered: tutored students gained 6.63 points more than ability-matched peers (p=.007), and the gain was concentrated in written reasoning, where the share of answers reaching relational quality rose from 8% to 49% in the tutored arm against 8% to 27% in the holdout. Modality did not matter for learning. The instructor's own final, on file for all 86 randomized students, shows the same direction (2.6 points of 100, with no difference on a pre-treatment midterm). Voice nearly doubled conversational interaction and cost 2.8 times as much to deliver, yet it produced weekly mastery statistically equivalent to text, even as students came to prefer it. Pedagogical structure shapes what students practice, and modality shapes how they interact with the tutor. Making an AI more humanlike does not by itself make it more educational.","authors":["Shihao Yang","Marshall Van Alstyne","Chrysanthos Dellarocas"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23958","pdf_url":"https://arxiv.org/pdf/2609.23958","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教育","随机实验","人机交互"],"reason":"研究AI导师教学效果，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.24934","version":1,"title":"Whose Facts Count? A Culturally Responsive Audit of LLM Evaluation Benchmarks","zh_title":"谁的事实算数？对LLM评估基准的文化响应性审计","abstract":"LLM benchmarks function as evaluation instruments, informing decisions that affect education, labor, and public services worldwide. Drawing on Hood, Kirkhart, and Hopson's culturally responsive evaluation (CRE) frameworks, this paper applies a six-dimension CR rubric to audit OpenAI's SimpleQA (N = 4,326 items) and the LMSYS Chatbot Arena (N = 600 conversations). Every SimpleQA question requires English-language archival verification as its evidentiary basis. A single rater's preoccupation with Colombian founding dates accounts for 2.70% of items, inflating the appearance of Global South coverage. English-language prompts constitute 76.3% of Arena conversations, against an International Telecommunication Union (ITU)-estimated 25.9% share of global internet users. A 50-item counter-benchmark scored a mean CR deficit nearly three times lower than SimpleQA (Cohen's d = 1.01). The paper proposes a practical CR evaluation framework. These are structural validity failures, not incidental measurement problems, with direct consequences for communities whose knowledge traditions these instruments were not built to see.","authors":["Fatima Tuz Zahra","Md. Sajeebul Islam Sk.","Rachel Chung"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24934","pdf_url":"https://arxiv.org/pdf/2609.24934","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","文化偏见","基准审计"],"reason":"论文审计LLM基准的文化响应性，属于纯NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.22694","version":1,"title":"Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review","zh_title":"面向痴呆相关车祸严重性预测的可审计与校准AI：支持人工复核的选择性延迟框架","abstract":"Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.","authors":["Gaurab Chhetri","Anika Baitullah","Subasish Das"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22694","pdf_url":"https://arxiv.org/pdf/2609.22694","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["交通安全","机器学习","选择性延迟"],"reason":"论文研究车祸严重性预测，属于交通安全领域，不涉及用LLM仿真人类被试或社会行为。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:13","error":null,"has_summary":false,"summary":null},{"id":"2609.23201","version":1,"title":"Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation","zh_title":"不要信任基准：通用LLM排名的局限性与任务特定评估的案例","abstract":"Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \\textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.","authors":["Danial Amin"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23201","pdf_url":"https://arxiv.org/pdf/2609.23201","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","基准测试","任务特定评估"],"reason":"论文讨论LLM基准测试的局限，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.23690","version":1,"title":"Inferring the microscopic mechanisms of opinion dynamics using a kinetic Ising model","zh_title":"用动力学伊辛模型推断意见动态的微观机制","abstract":"Kinetic Ising models are widely used to describe binary opinion dynamics, but their microscopic validity has rarely been tested empirically. Here, we infer the transition probabilities governing opinion updates from a year-long online social network and show that they are accurately described by an Ising heat-bath dynamics. The inferred parameters admit a direct sociological interpretation: the external field quantifies intrinsic bias, the coupling strength measures social influence, and a persistence term captures temporal inertia. We further show that persistence is positively correlated with node degree, while both persistence and interaction strength are strongly correlated with global network heterogeneity and clustering. Using the inferred time-dependent parameters in Monte Carlo simulations on the empirical temporal networks, we accurately reproduce the observed response functions, flip probabilities, and macroscopic opinion dynamics. Within the Twitter climate debate, our results provide direct empirical support for a kinetic Ising description of online opinion formation and establish a quantitative link between microscopic social behavior and evolving network topology.","authors":["Ixandra Achitouv","David Chavalarias","Vincent Lahoche"],"categories":["physics.soc-ph","cond-mat.stat-mech","cs.CY","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23690","pdf_url":"https://arxiv.org/pdf/2609.23690","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["意见动力学","伊辛模型","社会网络"],"reason":"论文研究人类舆论动力学，未使用LLM，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.24016","version":1,"title":"Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech","zh_title":"AI系统部署前的情境感知评估：尼日利亚金融科技的监管框架","abstract":"Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulatory instrument specifies what pre-deployment evaluation such systems must undergo before procurement. This paper reviews African fintech AI governance across global, continental, and Nigerian instruments, and shows that safety is affirmed as a principle while pre-deployment evaluation is operationally unspecified. Generic safety benchmarks cannot surface the failure modes most relevant to this domain, since none contain Nigerian institutional content or test for false positive misclassification of legitimate financial communications. These claims are demonstrated using SafeAlert, a purpose-built evaluation kit applied to six commercial models across three system prompt conditions. Results show that models resisting generic harmful content requests still produce complete fraud scripts under specific framing, and that several models misclassify most legitimate Nigerian bank communications as suspicious or fraudulent, a failure invisible to standard safety evaluation. The paper concludes with a regulatory framework proposing pre-deployment evaluation requirements for the CBN, NITDA, SEC, and the AU, arguing that the identified gap reflects an absence of regulatory specification, not a shortage of technical or financial resources.","authors":["Andrew Anogie Uduimoh","Hadiza Umar Yusuf","Oluwafemi Osho"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24016","pdf_url":"https://arxiv.org/pdf/2609.24016","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI安全评估","金融科技监管","模型评测"],"reason":"论文评估LLM在金融欺诈检测中的安全性，属于模型能力评测，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.24107","version":1,"title":"A Task-Oriented Multi-Agent Framework for Complex Wearable Health Analysis","zh_title":"面向复杂可穿戴健康分析的任务导向多智能体框架","abstract":"Wearable health questions often combine data retrieval, longitudinal analysis, and health advice over structured records. Prompting a single large language model with a complete record and a composite query obscures whether every request is executed and which evidence supports the answer. We propose a task-oriented multi-agent framework that represents a composite query as distinct intents and typed tasks with explicit intra-intent dependencies. Specialized agents execute retrieval, analysis, and advice tasks; isolated intent states preserve request boundaries and evidence relationships before aggregation. We evaluate the framework on a synthetic dataset of $10{,}000$ virtual users with one month of longitudinal wearable records, covering structured data retrieval, multi-intent recognition, and overall response quality. Across $1{,}500$ retrieval questions, the Query Agent achieves $98.3\\%$ accuracy, compared with $97.9\\%$ for the Direct LLM baseline, while reducing average query-stage token consumption from $6{,}869$ to $3{,}136$. On $180$ multi-intent questions, the Manager Agent achieves $100.0\\%$ Multi-Intent Coverage and $94.4\\%$ Multiset Jaccard Similarity. Under the current synthetic evaluation setting, our method receives higher mean Trustworthiness and Transparency scores on both question categories, whereas Actionability does not improve consistently. These results provide preliminary evidence that explicit task organization can support task-relevant data access and data-grounded longitudinal analysis, while leaving health advice generation and validation on real wearable data as open challenges.","authors":["Kunpeng Yang"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24107","pdf_url":"https://arxiv.org/pdf/2609.24107","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","可穿戴健康","任务分解"],"reason":"纯多智能体协作解决健康数据分析任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.22682","version":1,"title":"Self-Organizing Agent Teams Learn to Reason Together","zh_title":"自组织智能体团队学会共同推理","abstract":"Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $\\rho=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.","authors":["Aneesh Pappu","Mirac Suzgun","Yongchan Kwon","Federico Bianchi","Batu El","Mykel J. Kochenderfer","Hancheng Cao","James Zou"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22682","pdf_url":"https://arxiv.org/pdf/2609.22682","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协作推理","团队组织"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.23734","version":1,"title":"The Geometry of Alliances: Vote Transfer Modelling in French Two-Round Elections","zh_title":"联盟的几何学：法国两轮选举中的选票转移建模","abstract":"Two-round legislative elections are decided not only by first-round vote shares, but by how voters whose preferred party did not advance redistribute their votes among the surviving candidates. We develop a principled model of this transfer process, grounded in multi-dimensional ideological embeddings derived from the Chapel Hill Expert Survey, and apply it to French legislative elections. Calibrated on the 2017 and 2022 elections, the model is evaluated on the unusually complex 2024 contest, in which three ideologically distinct blocs reached the second round simultaneously, achieving 90.34 percent constituency-level accuracy. We show that ideological proximity alone explains the large majority of vote transfers, that a simple left-right axis is insufficient to capture the relevant distances, and that the structural configuration of the 2024 electorate was far less favorable to the far right than pre-election forecasts suggested.","authors":["Emmanuel Omont"],"categories":["cs.GT","physics.soc-ph"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23734","pdf_url":"https://arxiv.org/pdf/2609.23734","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["选举建模","选票转移","意识形态嵌入"],"reason":"论文研究选举中的选票转移建模，使用意识形态嵌入，未涉及LLM或人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2608.18265","version":4,"title":"Modeling Human Behavior with Type Vectors Using AI","zh_title":"使用AI类型向量建模人类行为","abstract":"We introduce a general, easy-to-implement AI-based modeling technique for analyzing human behavior. A key feature of this approach, which contrasts with existing modeling techniques, is that it combines the flexibility and interpretability of natural language with a mathematical structure that can be fitted to data and easily analyzed. We assign a large language model a vector of trait intensities-a type vector-and then ask it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) could correspond to \"You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5,\" after which it is asked to make choices. We can then vary the traits (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. We illustrate the method by applying it to model 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles. We find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. The type vectors needed to fit individuals across games cluster into fewer than a dozen groups, with substantial variation in fit across subjects. Moreover, the individual type vectors can predict behavior in held-out games with different rules and available actions. More broadly, this new modeling method is highly generalizable and interpretable: we can input any vector of traits and use them to model behavior across any setting","authors":["Matthew O. Jackson","Benjamin S. Manning","Yutong Xie","Walter Yuan","Qiaozhu Mei"],"categories":["econ.TH","cs.AI"],"primary_category":"econ.TH","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-08-20","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.18265","pdf_url":"https://arxiv.org/pdf/2608.18265","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为建模"],"reason":"用LLM模拟人类经济决策，并与大规模真实人类数据对照，直接命中核心方向。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":1,"question":"人类行为是否可以用少数几个特质维度（类型向量）来近似和预测？","design":"用大语言模型（LLM）作为仿真被试，通过提示词赋予其不同特质强度（如利他、风险厌恶、信任等）构成类型向量，然后让模型在10个经典经济博弈角色中做决策，通过调整类型向量最小化与人类选择的距离。","baseline":"119,147条决策，来自78,657名被试，覆盖35个国家，在10个经典经济博弈角色中的真实选择。","findings":"人类行为可以用三个维度（风险厌恶、策略复杂性、信任）紧密匹配；个体类型向量在博弈间聚类成少于12个群体，且能预测未参与拟合的博弈中的行为。","reliability":"论文未讨论","relevance":"直接命中核心方向：用LLM模拟人类经济决策并与大规模真实人类数据对照，方法可解释且可推广，值得精读原文。","inspiration":"借鉴其用可调类型向量作为LLM提示来系统扫描特质空间并拟合真实行为的方法，可迁移到资产定价实验中的风险偏好与信念异质性建模。｜设计一个研究：用LLM模拟投资者，赋予不同风险厌恶和过度自信水平，在实验性资产市场中交易，结果变量为价格泡沫程度和交易量，与真实实验市场数据（如Smith et al. 1988）对照，检验类型向量能否复现泡沫。"}},{"id":"2609.21636","version":1,"title":"Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30","zh_title":"引导大语言模型在挪威MFQ-30上向道德基础靠拢","abstract":"Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.","authors":["Hans Andersen","David Dichas"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21636","pdf_url":"https://arxiv.org/pdf/2609.21636","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德基础","算法保真度"],"reason":"用LLM复现人类道德基础分布，并与挪威样本对照，评估仿真可靠性与偏差，含批判性…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":4,"question":"LLM 在挪威版道德基础问卷（MFQ-30）上的道德画像是否稳定，能否通过提示或激活干预向目标人群靠拢？","design":"用六个开源 LLM 扮演挪威受访者，回答挪威 MFQ-30 问卷；施加两种处理：提示层 persona 引导（中性北欧受访者人设）和激活层 ActAdd 引导（对比向量注入）；结果变量为五个道德基础得分及与人类样本的马氏距离。","baseline":"Enstad 和 Finseraas (2024) 收集的 1282 名挪威受访者 MFQ-30 数据，已过滤不专注样本。","findings":"半数模型在注意力检查下真正作答，另一半输出扁平或趋中，平均接近人类但不追踪题目内容。中性北欧人设使作答模型与挪威均值的马氏距离缩小 44–77%，而 ActAdd 在固定中间层会扁平化画像而非定向引导单个基础。","reliability":"论文指出部分模型在基线时不作答，注意力检查可识别；人设引导可能诱发“认知幻影”，即模型表现出原本没有的作答行为，提示仿真结果可能不可靠。","relevance":"直接命中研究者对 LLM 仿真可靠性与偏差的关注，提供了真实人类对照和批判性发现，值得精读原文以了解注意力检查与引导干预的具体实现。","inspiration":"借鉴其注意力检查与第一 token 概率解码方法，可提高 LLM 问卷作答质量并识别无效仿真｜可迁移到经济政策偏好调查或消费者态度测量，如用 LLM 模拟不同人群对税收、福利或环保政策的道德评价｜用开源 LLM 扮演不同社会经济群体，施加中性人设或激活引导，测量其对政策陈述的同意程度，并与真实调查数据（如欧洲社会调查）对照，检验仿真偏差。"}},{"id":"2609.21259","version":1,"title":"CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition","zh_title":"CogGym：迈向人类与机器认知的大规模比较评估","abstract":"Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.","authors":["Lance Ying","Jinzhou Wu","Yingshan Susan Wang","Shivam Aarya","Luca M. Schulze Buschoff","Harry Chen","Katherine M. Collins","Andrea de Varda","Shuhao Fu","Sean Dae Houlihan","Akshay K. Jagadish","Guangyuan Jiang","Samuel Kiegeland","Tetsu Kurumisawa","Rongzhi Liu","Ryan Liu","Ningshan Ma","Kathryn McGregor","Younes Strittmatter","Polina Tsvilodub","Jacob Hoover Vigly","Sarah Wu","Enjie Xu","Yiling Yun","Kelsey Allen","Tyler Brooke-Wilson","Brian Christian","Evelina Fedorenko","Michael C. Frank","Michael Franke","Tao Gao","Samuel J. Gershman","Robert D. Hawkins","Jennifer Hu","Julian Jara-Ettinger","Max Kleiman-Weiner","Sydney Levine","Tal Linzen","Hongjing Lu","Timothy O'Donnell","Desmond C. Ong","Steven T. Piantadosi","Rebecca Saxe","Eric Schulz","Tianmin Shu","Felix A. Sosa","Ilia Sucholutsky","Tan Zhi-Xuan","Tomer Ullman","Fei Xu","Ilker Yildirim","Jian-Qiao Zhu","Thomas L. Griffiths","Tobias Gerstenberg","Kevin Smith","Joshua B. Tenenbaum"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21259","pdf_url":"https://arxiv.org/pdf/2609.21259","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知实验","人类对照"],"reason":"直接比较LLM与人类在认知实验中的行为，有真实人类数据对照，并评估模型-人类拟…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":3,"question":"如何在大规模、多样化的认知实验中系统比较AI模型与人类的行为，以识别模型与人类判断的相似与分歧？","design":"构建CogGym框架，将258个来自100篇论文的人类常识推理认知实验标准化为统一的实验标记语言（EML），并评估50个大型语言模型在这些实验上的表现，测量模型输出与人类反应分布的拟合度。","baseline":"原始认知实验中的真实人类被试反应数据，包括文本、图像和视频模态，并报告人类分半信度作为上限。","findings":"模型规模越大、越新，与人类判断的拟合度越高，但在常识推理任务上的进步速度慢于数学和编程等正式推理基准。最佳模型与人类的拟合度（文本R²=0.59，图像R²=0.58，视频R²=0.43）仍远低于人类分半信度（文本R²=0.93，图像R²=0.95，视频R²=0.92）。","reliability":"论文指出模型-人类拟合度远低于人类分半信度，表明模型尚未完全捕捉人类认知的细微结构；同时，标准化过程中可能存在实验转换的失真，且当前覆盖范围限于常识推理，未涉及其他认知领域。","relevance":"该研究直接比较LLM与人类在大量认知实验中的行为，有真实人类数据对照，并评估模型-人类拟合度，与研究者关注的人类仿真实验高度相关，值得精读以了解大规模评估框架和模型偏差。","inspiration":"借鉴其半自动化实验标准化流程和分半信度作为上限的评估方法，可迁移到经济决策实验（如风险偏好、时间贴现、博弈行为）的仿真验证。｜可应用于消费者跨期选择或资产定价实验，检验LLM是否复现人类的时间不一致性或风险厌恶。｜设计：以LLM为被试，施加跨期选择任务（如现在100元 vs. 一个月后120元），测量贴现率，并与真实人类实验数据（如Andersen et al., 2008）对照，计算模型-人类拟合度与人类分半信度的差距。"}},{"id":"2609.20827","version":1,"title":"From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators","zh_title":"从出院记录到患者理解：基于人格的开放式模拟将LLM作为出院教育者","abstract":"Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.","authors":["Won Seok Jang","Zonghai Yao","Hong Yu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20827","pdf_url":"https://arxiv.org/pdf/2609.20827","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","患者教育","人类数据对照"],"reason":"用LLM模拟患者进行出院教育评估，有真实临床数据对照，涉及医疗场景，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":4,"question":"如何评估大语言模型在开放式、多轮对话中作为出院教育者，使不同人格与健康素养的患者真正理解出院指导的能力？","design":"构建 DischargeBench 仿真框架：用 LLM 扮演虚拟患者（基于 MIMIC-IV 真实出院记录，设定人格、教育水平、健康素养、病史回忆等属性），候选 LLM 扮演教育者进行多轮对话，并由教育监控代理仅调节患者侧以保持真实性；通过 LLM 裁判对对话质量、主题覆盖、理解程度和事实一致性四个维度评分。","baseline":"使用 MIMIC-IV-Ext-DischargeBench 的 477 个真实病例（来自 MIMIC-IV 和 MIMIC-IV-Note），并由医生标注用于对齐 LLM 裁判的评分。","findings":"GPT-5 系列在对话质量和主题覆盖上领先，但可读性较差；总体分数掩盖了不同 ICD 章节和患者人格下的显著差异，困难人格暴露了覆盖失败、理解差距和源答案一致性下降。","reliability":"论文指出 LLM 评估应关注患者理解而非文本质量或答案准确性；困难人格下仿真会失效，且 LLM 裁判的评分可能与医生标注存在偏差。","relevance":"该研究用 LLM 模拟患者进行出院教育评估，有真实临床数据对照，并批判性指出仿真在困难人格下失效，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其多智能体仿真设计：用 LLM 扮演异质性个体并设置监控代理防止污染处理信号，同时用真实数据校准裁判评分。｜可迁移到消费者金融教育或政策沟通场景，如评估 AI 理财顾问对不同金融素养人群的讲解效果。｜设计：用 LLM 模拟不同金融素养和人格的投资者作为被试，处理为 AI 顾问的个性化解释，结果变量为投资决策质量和风险理解，对照真实投资者调查数据（如 FINRA 金融素养调查）。"}},{"id":"2609.21439","version":1,"title":"People escalate against a competitor labelled human and hold back against one labelled an optimising machine","zh_title":"人们面对标记为人类的竞争者会升级投入，面对标记为优化机器的竞争者则会退缩","abstract":"People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1,395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral.","authors":["Vinicius Ferraz","Leon Houf"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21439","pdf_url":"https://arxiv.org/pdf/2609.21439","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","行为实验","人机交互"],"reason":"用LLM作为对手与人类被试互动，研究标签信息对行为的影响，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":11,"question":"在人类与AI竞争的情境中，标签信息（告知对手是人还是AI）是否独立于对手实际行为影响人们的竞争升级行为？","design":"采用预注册实验（N=1395），使用动态全支付拍卖（重复消耗战）作为竞争任务。通过欺骗自由设计交叉两个维度：告知信息（对手可能是人类、模仿人类的AI、优化竞争的AI或无信息）与实际对手（人类、模仿人类的AI、优化竞争的AI）。测量结果变量为出价升级（中位数价格）和双方均获正收益的概率。","baseline":"真实人类被试（Prolific平台招募）与真实人类对手、两种AI对手（模仿人类和优化竞争）的实际对局数据。","findings":"信息效应显著：仅告知对手可能是人类使中位数出价上升6.7点，告知可能是优化机器使中位数出价下降8.8点，差距约为奖品价值的15%。对手效应表现为与AI实际竞争降低出价，但同时减少双方均获正收益的概率约三分之二。","reliability":"论文未讨论","relevance":"该研究用真实人类被试与AI对手互动，通过标签信息操纵考察行为变化，有真实人类数据对照，直接回应了LLM仿真中信息环境对行为的影响，值得精读以理解信息效应与对手效应的分离方法。","inspiration":"该研究采用欺骗自由设计交叉告知信息与实际对手，分离信息效应与对手效应，方法上可借鉴用于经济实验中标签或框架效应的因果识别。｜可迁移到算法定价或自动竞价场景，研究披露算法身份对市场参与者竞争行为的影响。｜设计实验：招募人类被试参与模拟拍卖或议价博弈，随机告知对手为人类或算法（实际对手固定为算法或人类），测量出价或报价行为，并与真实市场交易数据（如电商平台竞价记录）对照。"}},{"id":"2609.16374","version":2,"title":"When a Story Feels Like Mine: How Personalized Narratives and Humor Shape Older Adults' Empathy toward LLM-Generated Peer Health Stories","zh_title":"当故事感觉像我的：个性化叙事与幽默如何塑造老年人对LLM生成同伴健康故事的同理心","abstract":"Peer stories have been shown to boost self-efficacy in older adults' health behavior change. Despite their effectiveness, peer stories are difficult to deploy in health promotion at scale given the difficulty of matching the diverse health concerns and coping styles of heterogeneous older populations. Large language models (LLMs) have been shown to generate authentic narratives, yet how personalization and narrative affective style, such as humor, jointly shape older adults' responses remains unknown. We developed a theory-driven system that generates first-person peer health narratives varying in personalization and humor through a three-stage LLM pipeline grounded in self-efficacy mechanisms. Thirty-one older adults were invited to participate in a within-subjects lab study. Results showed that personalization increased perceived relatability and relevance of peer stories, especially for older adults with lower humor preference. These findings position individual differences in affective styles as a second dimension in designing personalization for LLM-assisted health communication.","authors":["Kexin Quan","Precious Olalere","Smit Desai","Jessie Chin"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-16","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.16374","pdf_url":"https://arxiv.org/pdf/2609.16374","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM生成叙事","个性化健康传播","老年人用户研究"],"reason":"用LLM生成健康叙事并测量老年人反应，有真实人类被试对照，但非直接仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":7,"question":"个性化与幽默如何共同影响老年人对LLM生成的同伴健康故事的共情及相关感知？","design":"本研究并非直接仿真人类被试，而是用LLM（gpt-4o-mini）生成第一人称同伴健康叙事，通过三阶段流水线实现个性化（基于参与者自报的健康挑战、目标、障碍和应对策略）和幽默（亲和、应对导向）的操纵。31名社区老年人参与2（个性化）×2（幽默）的被试内实验，评估对故事的共情、理解、相关性、真实性和偏好。","baseline":"无对照（未使用真实人类数据作为基准，而是直接测量人类被试对LLM生成内容的反应）。","findings":"个性化显著提高了老年人对故事的感知相关性和共鸣，尤其对幽默偏好较低的个体效果更强。幽默本身并未增强共情或相关性，但可能增加故事的吸引力。","reliability":"论文未讨论（节选内容未提及失效条件或局限）。","relevance":"该研究虽非直接仿真人类被试，但提供了LLM生成个性化叙事并测量人类反应的实验范式，对关注LLM在健康传播中应用的仿真研究者有参考价值，但缺乏真实人类数据对照，与经济学实验仿真关联较弱。","inspiration":"借鉴其个性化处理的设计：基于个体特征（如健康目标、障碍）动态生成内容，并操纵叙事风格（幽默）作为第二维度，采用被试内设计测量感知结果。｜可迁移到消费者金融决策中的个性化信息干预，如退休储蓄建议或健康保险选择，研究不同叙事风格对决策的影响。｜以真实投资者为被试，用LLM生成个性化投资建议（基于其风险偏好和财务目标），操纵叙事语气（幽默 vs. 严肃），测量投资意愿和风险感知，并与历史投资行为数据对照。"}},{"id":"2609.20846","version":1,"title":"Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models","zh_title":"奖励高效推理可改善推理模型在欠明确任务上的弃权行为","abstract":"While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models' efficiency (44% shorter CoTs on average).","authors":["Polina Tsvilodub","Max H\\\"oth","Michael Franke","Bj\\\"orn Deiseroth","Carina Kauf"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20846","pdf_url":"https://arxiv.org/pdf/2609.20846","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM弃权行为","人类对照","可靠性评估"],"reason":"用人类数据对照，评估LLM在不确定任务上的弃权行为，并改进其与人类一致性，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":8,"question":"大型推理模型（LRMs）在任务信息不足时是否无法像人类一样高效地弃权，以及如何通过奖励设计改进其弃权行为。","design":"该研究并非严格意义上的仿真实验，而是将多个LRMs（4B–32B，六个模型家族）在QuestBench和AbstentionBench上的弃权表现与人类行为进行对比，然后提出SURE奖励（结合结果奖励与过程奖励）对4B模型进行GRPO微调，测量弃权准确率和思维链长度。","baseline":"研究者自行开展了一项人类实验，让人类被试对可答与不可答问题进行二元判断，记录其准确率和推理努力（以回答时间或步数衡量），作为LRMs的对照基准。","findings":"LRMs在不可答任务上生成的思维链比可答任务更长，弃权准确率远低于人类；使用SURE奖励微调后，模型弃权准确率平均提升12.8%，思维链长度平均缩短44%，同时保持回答能力。","reliability":"论文承认SURE的过程奖励主要针对信息缺失类弃权，对于安全等其他弃权原因可能需要不同的过程监督；微调模型仅在有限数据集上评估，需在更多弃权数据集上验证；人类实验中被试在不可答问题上有时会基于特定假设作答，与模型的假设差异需进一步比较。","relevance":"该研究直接对比LLM与人类在不确定任务上的弃权行为，并利用人类数据改进模型，属于用LLM仿真人类决策并评估可靠性的工作，对关注经济学实验和政策评估中模型弃权行为的研究者有参考价值。","inspiration":"借鉴其将人类行为作为基准并设计奖励函数对齐人类效率的做法，可用于经济决策仿真中的理性疏忽或信息获取行为。｜可迁移到消费者在信息不完全时的购买决策、投资者面对模糊信息的资产配置、或政策公告下预期形成等场景。｜以LLM作为被试，设计信息缺失的经济决策任务（如模糊收益彩票），处理为是否提供额外信息或奖励函数中惩罚过度推理，结果变量为弃权率与推理步数，对照真实人类实验数据（如实验室彩票选择实验）。"}},{"id":"2609.21401","version":1,"title":"Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue","zh_title":"与机器对话的错位：人机对话中的道德、礼貌与对齐","abstract":"Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are evaluated, trusted, and designed. This study investigates how morality, politeness, and alignment - three dimensions central to cooperative dialogue - function in human-AI interaction compared to human-human conversation. We analyze 15,881 human-ChatGPT and 10,784 human-human multi-turn dialogues, using mixed-effects models to identify which features predict turn-to-turn alignment. We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture. Moral output appears preconfigured rather than negotiated; warmth is generated without face sensitivity; linguistic convergence declines persistently. Most strikingly, the cooperative mechanisms themselves reverse direction: hedging and softening associated with greater accommodation between humans are associated with reduced alignment when produced by AI, and purity framing associated with human divergence coincides with users converging toward the AI. Agency - giving users room to shape the exchange - is the most consistent predictor of alignment across both interaction types, while lower moral assertiveness in more recent models is not accompanied by better cooperation. Together these patterns suggest that AI reproduces the surface of cooperation without the mutual adaptation that grounds it between humans - and, more surprisingly, that mechanisms sustaining human accommodation can run in reverse with AI, suggesting a turn-level view may be insufficient for interaction-level success.","authors":["Marina Mitiaeva","Lu Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21401","pdf_url":"https://arxiv.org/pdf/2609.21401","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人机对话","合作沟通","仿真偏差"],"reason":"用LLM与人类对话数据对照，分析合作沟通机制，揭示AI表面模仿但缺乏真实社会适…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":21,"question":"在人类与AI对话中，道德、礼貌与对齐三个合作沟通维度如何运作，与人类间对话相比有何差异？","design":"本研究不是仿真实验，而是对真实对话语料的计算分析：使用WildChat中15881条人类与ChatGPT多轮对话和10784条人类间多轮对话，用混合效应模型分析道德表达（MFT分数）、礼貌标记（32个标记）和逐轮对齐（词汇、句法、语用、情感四个通道的相似度）之间的关系。","baseline":"人类间对话数据（10784条多轮对话）作为对照基准。","findings":"AI产生合作沟通的表面特征但缺乏底层社会架构：道德输出是预配置而非协商的，温暖生成没有面子敏感性，语言趋同持续下降。合作机制方向反转：人类中与更大适应相关的模糊限制语和软化在AI产生时与对齐降低相关，纯洁框架在人类中与分歧相关但用户却向AI趋同。","reliability":"论文未讨论仿真失效条件，但指出AI的合作机制可能反向运行，提示基于回合的视角可能不足以评估交互层面的成功。","relevance":"该研究直接对比人类与AI对话中的合作机制，揭示AI表面模仿但缺乏真实社会适应，对理解LLM在交互中的行为偏差和可靠性有重要价值，值得阅读原文。","inspiration":"借鉴其大规模语料对照分析和多维度测量方法，可迁移到经济金融中的沟通与信任场景，如投资者与AI顾问的对话。｜设计一个实验：让人类被试与LLM扮演的金融顾问进行投资建议对话，测量道德语言、礼貌标记和语言对齐，并与人类顾问对话对照，结果变量为投资决策和信任度。"}},{"id":"2609.21149","version":1,"title":"Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake","zh_title":"面向AI辅助精神病学接诊的临床医生导向质量保证","abstract":"Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.","authors":["King Shi","Amanda Li","Jonathan Ivey","Synthia Qia Wang","Guan Gui","Hyunseo Kim","Peter Zandi","Jason Straub","Jacob Taylor","Ananya Joshi"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21149","pdf_url":"https://arxiv.org/pdf/2609.21149","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM患者仿真","临床质量保证","人机对照"],"reason":"用LLM模拟患者进行精神病学访谈，并与临床医生对照，属于人类仿真且有人类数据基…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":20,"question":"如何为AI辅助精神科接诊系统建立一个以临床医生为基准、可重复且低负担的质量保证评估平台？","design":"使用记忆增强的患者模拟器InterviewPlayground，基于专家编写的病例摘要创建交互式合成患者；构建模拟接诊平台，让6名临床医生和一个基于GPT的LLM接诊员分别对合成患者进行25分钟访谈；通过回忆表单和转录分析测量信息恢复率、无依据推断比例和安全问题特征化率。","baseline":"6名执业临床医生在相同合成患者上的访谈表现作为人类基准。","findings":"LLM接诊员恢复了更多临床相关条目（88.0%对38.9%），但做出了更多无依据的临床推断（56.8%对27.8%），且更少特征化已识别的安全问题（33.3%对66.7%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟患者进行精神病学访谈，并与临床医生对照，属于人类仿真且有人类数据基准，但场景是医疗质量保证而非经济学实验，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"借鉴其使用合成患者固定临床真相以支持跨风格比较和重复评估的方法，可迁移到消费者金融决策或信贷审批等需要标准化对话场景的经济学实验中，例如用LLM模拟不同特征的借款人，让人类信贷员和AI审批系统分别进行访谈，测量信息提取准确性和决策偏差，并与真实信贷数据对照。"}},{"id":"2608.25180","version":2,"title":"Self-Explanation Tutor for Active Study of CS1 Worked Examples","zh_title":"用于主动学习CS1工作示例的自解释导师系统","abstract":"Worked examples are an important part of introductory programming, but reading their expert explanations is passive. Self explanation, students explaining the problem and its solution to themselves with subgoal level analysis, converts passive reading into an active study of worked example, yet it is hard to scale because assessing free-text explanations and returning timely feedback has had no easy automated solution. We investigate whether a large language model (LLM) can fill that gap. We build a self-explanation tutor for introductory programming, ESSE, in which students explain lines of worked examples and receive immediate LLM feedback on the correctness and completeness of each explanation, and we pursue two goals. First, we ask whether the LLM judges student explanations well enough to serve as the engine of the tutor; we assess its judgments against two independent human reference standards of different kinds, a single domain expert and a crowd of non-expert raters, each with its own strengths and weaknesses, characterizing both where the LLM is reliable and the systematic tendencies in how it diverges. Second, we ask whether the LLM-based tutoring benefits students; deploying it in an introductory Java course, we find that its feedback leads students to persist and revise rather than abandon a line, that their explanations grow more complete and conceptually richer across attempts, and that students show evidence of learning. These indicate that LLM-based assessment is good enough to power a self-explanation tutor, and that the tutor positively shapes how students study worked examples.","authors":["Arun-Balajiee Lekshmi-Narayanan","Mohammad Hassany","Kamil Akhuseyinoglu","Rully Hendrawan","Peter Brusilovsky"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-08-27","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.25180","pdf_url":"https://arxiv.org/pdf/2608.25180","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","教育技术","人类判断对照"],"reason":"LLM替代人工评估学生解释，属标注替代而非仿真人类被试，但涉及与人类判断对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2609.19150","version":2,"title":"Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations","zh_title":"采样揭示风格：在LLM激活中无监督、免训练地发现提示条件风格轴","abstract":"Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis (PCA) to the pooled hidden activations, and label the resulting axes automatically from the pole generations. We validate the discovered axes against 245 human-elicited stylistic annotations in a two-phase study. On our strongest model (Qwen-3.5-4B-Instruct), the top two axes match spontaneously requested human dimensions with 72.8% precision and 43.6% macro-recall, and 75.6% of validity ratings judge the axes' polar generations accurate to their labels, with 90.9% adjacent inter-annotator agreement. Discoverability is strongly model-dependent: both Qwen models and Llama-3.2-3B expose human-salient axes, while DeepSeek-7B-Chat drops to 35.3% precision, its leading components dominated by structural rather than stylistic variance. Simple PCA over a model's own decoding variance is thus an effective, low-cost probe of stylistic structure in LLM representations, one that also exposes sharp cross-model differences in how that structure is organized.","authors":["Ajit Mallavarapu","Ziwei Gu"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-18","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.19150","pdf_url":"https://arxiv.org/pdf/2609.19150","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM表征","风格分析","可解释性"],"reason":"研究LLM风格表征，用人类标注验证，但测的是模型而非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.20829","version":1,"title":"SAGE: Schema-Guided LLMs for Grant Review","zh_title":"SAGE：模式引导的LLM用于资助评审","abstract":"Grant reviewers must apply detailed criteria to application forms, budgets, and supporting documents while producing assessments that colleagues can inspect. We present SAGE, Schema-Guided Aspect-Based Grant Evaluation, a system that translates a grant rubric into structured checks and links its judgements to evidence from the application package. We evaluate SAGE in two stages on 35 nonprofit grant applications. A post-factum comparison with 105 reviews from the original competition shows fair ordinal agreement (kappa = 0.29). The foundation then conducted a criterion-level re-review after inspecting SAGE, producing 202 assessments. In this assisted round, SAGE reached kappa = 0.58 and outperformed a one-prompt-per-criterion baseline (kappa = 0.33 on the common subset), with higher rank correlation and lower error. A claim-level audit further identifies confirmed, disputed, and unaddressed parts of the structured draft. SAGE operationalizes the review methodology by producing a detailed, evidence-linked, and auditable draft for expert correction.","authors":["Erik Varapaev","Andrei Chetvergov","Stepan Ukolov","Timofei Sivoraksha","Alexander Evseev","Sergey Bolovtsov"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20829","pdf_url":"https://arxiv.org/pdf/2609.20829","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评审","资助申请","证据链接"],"reason":"LLM替代人工评审，非仿真人类被试，但涉及人类评审数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.21075","version":1,"title":"Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation","zh_title":"与生活经验对齐：心理健康支持生成中微调的异质性益处","abstract":"As access to professional mental healthcare remains limited, many individuals turn to online platforms such as Reddit to seek peer support situated within human lived experience. However, a significant portion of such queries go unanswered, presenting an opportunity for using Large Language Models (LLMs) to fill this gap. While LLMs have demonstrated strong performance on clinical benchmarks, their ability to generate lived-experience informed and community-aligned peer support is underexplored. Addressing this gap, we introduce the COmmunity-centered Peer Engaged Support (COPES) dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives to mental health support seeking queries. Evaluating zero-shot and post-trained (SFT and DPO) models, we show that post-training on COPES significantly improves Strategy Alignment (>50% for general-purpose models) and alignment in Emotion & Tone. However, we also observe that such improvements are heterogeneous and alignment improvements vary significantly across subreddits and requested coping strategies. Furthermore, post-training induces distributional shifts, heavily favoring problem-focused recommendations while suppressing emotion-focused strategies. Together, this work shows that while curating community-driven data improves the alignment of LLM responses, model performance remains disparate across distinct sub-communities and specific mental health needs.","authors":["Mohit Chandra","Nabin Kim","Eli Min","Aamogh Sawant","Tanmay Sutar","Munmun De Choudhury"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21075","pdf_url":"https://arxiv.org/pdf/2609.21075","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM对齐","心理健康支持","社区数据"],"reason":"评估LLM生成的心理支持回复与社区观点的对齐，属于测量模型输出而非仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.21094","version":1,"title":"Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models","zh_title":"价值观几何：语言模型中伦理偏好对齐的任务向量组合","abstract":"Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.","authors":["Utkarsh Agarwal","Monojit Choudhury"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21094","pdf_url":"https://arxiv.org/pdf/2609.21094","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["价值观对齐","任务向量","跨语言偏见"],"reason":"论文测量LLM的价值观偏好，属于把LLM本身当测量对象，而非用LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.21154","version":1,"title":"CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop","zh_title":"CoLearn：在人机共同学习循环中学习学习者的智能导师","abstract":"Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer as a single bit of signal. We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner's mastery and misconceptions. This memory is updated as evidence accumulates and is used to generate the next personalised question. CoLearn has three components: (i) a persistent learner-state memory that updates per-topic mastery with a soft-evidence variant of Bayesian Knowledge Tracing, where a large language model acts as a continuous observation function; (ii) adaptive question generation that targets the learner's weakest topic and recurring misconceptions; and (iii) an evidence view that makes personalisation visible and testable through live progress visualisation and blind A/B comparison. In blind A/B evaluation, questions conditioned on this memory are preferred over non-personalised ones 68-69% of the time, and in persona simulations with hidden ground-truth mastery the agent's belief converges toward the learner's true mastery.","authors":["Kailai He","Zhihao Wu","Linhai Zhang","Runcong Zhao","Yulan He","Jiazheng Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21154","pdf_url":"https://arxiv.org/pdf/2609.21154","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["智能辅导系统","学习者建模","LLM标注替代"],"reason":"LLM作为观察函数更新学习者模型，非仿真人类被试，但涉及用LLM替代人工标注，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.21857","version":1,"title":"Do Personality-Tuned LLMs Make Better Social Agents?","zh_title":"人格调优的大语言模型能成为更好的社会智能体吗？","abstract":"LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work investigates whether personality-aware fine-tuning can reduce this gap by improving the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone. We fine-tune two small open-weight LLMs, Qwen2.5-7B-Instruct and Ministral-8B-Instruct, using a corpus that combines personality-labelled social media posts and dialogues to create a personality-based dialogue engine for social simulation. The resulting models are evaluated across multiple social interaction scenarios using three independent LLM judges, which assess personality fidelity and provide evidence-based behavioral interpretations. We additionally quantify inter-rater agreement and lexical characteristics of the generated dialogue. Results indicate that fine-tuned models are not better at role-playing different personalities than their respective baseline models. However, low inter-rater agreement limits the confidence with which these results can be interpreted. Concerning the quality of generated texts, fine-tuned models are mostly comparable to the baselines, with fine-tuning improving the linguistic diversity of the Qwen models. While the results appear generally usable and the baseline models offer the best overall performance, future studies should place greater emphasis on the quality and domain alignment of training data for accurate personality role-playing.","authors":["Tim Krabbe","Xiaodan Shi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21857","pdf_url":"https://arxiv.org/pdf/2609.21857","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人格调优","角色扮演","社会模拟"],"reason":"研究人格调优LLM的角色扮演一致性，属于人格测量与对话生成，无真实人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2609.21626","version":1,"title":"One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction","zh_title":"一个提示并不适合所有人：面向个性化信息抽取的自元进化","abstract":"Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successful editing patterns. To enable scalable training and evaluation, we release a persona-driven IE benchmark of 292 simulated enterprise users, paired with a reproducible persona-generation pipeline grounded in O*NET occupational taxonomies. On this benchmark, Self-Meta-Evolve achieves a 74.58% success rate, outperforming the strongest prompt-optimization baseline by 13.56 absolute points, and reaches 52.54\\% within only two iterations. A double-blind human study with twenty real professionals further confirms that prompts adapted by our framework win against static baselines in 71% of pairwise comparisons.","authors":["Hongliang Li","Lu Wang","Yong Xu","Hanyang Chen","Zhitao Hou","Xiaoting Qin","Song Ge","Qingwei Lin","Dongmei Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21626","pdf_url":"https://arxiv.org/pdf/2609.21626","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["提示优化","个性化信息抽取","用户模拟"],"reason":"用LLM模拟企业用户进行信息抽取，属于替代人工标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.21841","version":1,"title":"EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise","zh_title":"EnterpriseVal：量化企业生成式AI的效能、可靠性与价值","abstract":"Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer \"what can the model do?\", whereas a deployment decision requires \"is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?\". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation","authors":["Abbas Raza Ali","Muhammad Ajmal Siddiqui","Moona Zahid"],"categories":["cs.AI","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21841","pdf_url":"https://arxiv.org/pdf/2609.21841","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","企业应用","预测驱动推断"],"reason":"用LLM替代人工评审，属标注替代而非仿真人类被试，但方法可借鉴","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.22067","version":1,"title":"Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw","zh_title":"日常AI代理使用中的价值敏感委托：来自OpenClaw的证据","abstract":"Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect's corpus share, values clustered not at the agent's outputs but at the operating conditions users set around a run. Values were usually met where users described what the agent delivered, in five of six groups, and mostly unmet where users described supervising it, in all six groups. We conceptualize this pattern as value-sensitive delegation. Supporting human values requires attention not only to what an agent accomplishes, but to the conditions users set around delegation, including cost, access, and oversight.","authors":["Renkai Ma","Ruyuan Wan","Xuan Lu","Fan Yang","Chen Chen","Lingyao Li"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22067","pdf_url":"https://arxiv.org/pdf/2609.22067","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["价值敏感设计","AI代理使用","LLM辅助分析"],"reason":"用LLM辅助分析人类使用AI代理的价值观，非仿真人类被试，但涉及LLM辅助分析…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.20942","version":1,"title":"When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation","zh_title":"当AI评审训练AI评审员：科学判断的崩塌与缓解","abstract":"Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018--2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern $\\textbf{scientific-judgment collapse}$. To mitigate this failure mode, we introduce $\\textbf{TrustReviewer}$, an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.","authors":["Sy-Tuyen Ho","Minghui Liu","Furong Huang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20942","pdf_url":"https://arxiv.org/pdf/2609.20942","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI审稿","模型训练","判断多样性"],"reason":"LLM替代人工审稿，属标注替代而非仿真人类被试，但涉及AI评价偏差，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.14819","version":2,"title":"A primer on evaluation methods for large language models in healthcare","zh_title":"医疗领域大语言模型评估方法入门","abstract":"Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.","authors":["Suzannah E McKinney","Phuc Vu","Samuel A Justice","Christopher Humphries","Alyssa Pradhan","Timothy J Keyes","Sarah F Mercaldo","James M Hillis"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-15","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.14819","pdf_url":"https://arxiv.org/pdf/2609.14819","source_feed":"cs.CL","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","医疗AI","综述"],"reason":"综述LLM在医疗中的评估方法，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.20902","version":1,"title":"Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes","zh_title":"用于动机性访谈的生成式人工智能聊天机器人：从系统设计到干预结果的范畴综述","abstract":"Motivational interviewing (MI) is a collaborative approach to elicit autonomous motivation for health behavior change. Generative AI (GenAI) offers new ways to deliver MI via conversational systems, but evidence on their design, assessment, and translation into interventions remains fragmented. This scoping review characterized evidence on GenAI-MI chatbots across system design, safety, MI quality, user perceptions, and intervention outcomes. We conducted a PRISMA-ScR scoping review. Nine datasets were searched for studies published or publicly available from January 1, 2015 to June 2, 2026 that used GenAI to generate MI chatbot responses or counselor utterances. Data were extracted using a predefined framework and synthesized descriptively. Forty-seven reports (48 studies) were included. Twenty (41.7%) focused on system design without direct participant use; 28 (58.3%) involved direct interaction. Most systems were text based and disembodied; 23 (47.9%) incorporated dynamic adaptation. Safety measures were unevenly reported. Among studies with direct use, 21/28 (75.0%) reported informed consent or user education. Thirty (62.5%) assessed MI quality, generally suggesting MI-consistent interactions. User perceptions were favorable, especially empathy, usability, helpfulness, and intention to use, though measures were heterogeneous. Eighteen (37.5%) reported intervention outcomes, mostly after a single session. Positive findings were more consistent for short-term motivation than sustained behavioral or functional change. GenAI-MI chatbots can deliver MI-consistent interactions perceived favorably, but evidence for sustained behavioral or functional change is limited. Future research should strengthen runtime safety monitoring, standardize MI quality assessment, and use longer-term comparative designs with behavioral and functional outcomes.","authors":["Runze Hu","Jingqi Kong","Yang Yang","Yihang Yang","Jingyao Liu","Haizhou Tang","Shanghang Zhang","Zheng Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20902","pdf_url":"https://arxiv.org/pdf/2609.20902","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["动机性访谈","聊天机器人","健康干预"],"reason":"该研究是GenAI聊天机器人用于动机性访谈，属于角色扮演对话，无实验或测量目的…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.21349","version":1,"title":"From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers","zh_title":"从记忆到行为：面向社交媒体影响者的行为感知角色扮演框架","abstract":"Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging. Existing in-context learning-based methods fail to capture how individuals react under different situations. In addition, LLM-based evaluation is difficult for obscure individuals. To address these challenges, we propose Situation--Internal state--Behavior Persona method to incorporate situation-dependent behavioral strategies. We further design an evaluation protocol that provides LLM evaluators with references about the impersonated individual. We evaluate our approach on a newly constructed dataset for the task of generating replies on social media. Experimental results show that our proposed method outperforms state-of-the-art ICL-based baselines, while our evaluation protocol achieves moderate correlation with human judgment. Besides, experiments on fictional-character benchmarks demonstrate that our proposed method is applicable beyond the social media setting. These findings suggest that incorporating behavioral information broadly improves the fidelity of role-playing for real individuals on social media or fictional characters.","authors":["Ji-Lun Peng","Yi-Zhen Zhang","Chun-Nan Chou","Yun-Nung Chen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21349","pdf_url":"https://arxiv.org/pdf/2609.21349","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","社交媒体","行为策略"],"reason":"角色扮演生成社交媒体回复，无实验或测量目的，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.00285","version":2,"title":"Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct","zh_title":"十六个模型，不到两种声音：在无唯一正确答案时测量集成离散度","abstract":"Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles place more than one reading before a decision-maker on the premise that several models supply several perspectives. Dispersion over their outputs is measured both as diversity and as uncertainty, and both traditions validate it against a correctness criterion that this task does not admit. Measuring diversity is a solved problem: the Vendi Score, the exponential of the von Neumann entropy of a similarity matrix, is an effective number of distinct elements. What a single aggregate does not say is where the diversity comes from. We define a per-model dissent contribution, the complement of a model's mean similarity to the other members of its ensemble: a magnitude from the same matrix, not a decomposition of the spectral index, whose maximum identifies the most divergent voice. Crossing model and case, we test as a preregistered hypothesis whether model identity accounts for a non-zero share of the variance in dissent, and characterise the structure that test detects. The panel formulated fifteen stratified vignettes, yielding 7,082 formulations for analysis. Model identity was a detectable structuring factor of the dissent that remained, but the usual categories recovered it only partly: scale differences pointed in opposite directions across pairs, family grouped models on only five two-member lines, and the most divergent voice changed with panel composition, so that the surfaced outlier describes the ensemble rather than the model. Dissent did not track the interpretive openness for which the case bank was stratified; it was organised by clinical content instead, leaving the dispersion an ensemble produces a property to measure rather than assume.","authors":["Mario Vega-Barbas","Lidia Mora-Valenciano","Iv\\'an Pau","Fernando Seoane","Farhad Abtahi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-08-04","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.00285","pdf_url":"https://arxiv.org/pdf/2608.00285","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多模型集成","输出多样性","心理治疗案例"],"reason":"研究多模型集成输出的多样性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.13430","version":2,"title":"Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity","zh_title":"你确定你确定吗？指令微调对置信度和词汇多样性的影响","abstract":"Instruction-tuned language models achieve strong performance across a range of generation tasks but have recently been shown to exhibit verbalized overconfidence, which may manifest in less diverse supporting rationales for incorrect answers. However, whether such overconfidence is associated with rationale consistency remains an open question. In this paper, we study whether changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently increases answer confidence, despite limited changes in predictive accuracy, while degrading likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.","authors":["Irina Proskurina","Mayank Kumar","Oyindolapo O. Komolafe"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-08-15","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.13430","pdf_url":"https://arxiv.org/pdf/2608.13430","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["指令微调","置信度校准","模型行为分析"],"reason":"研究指令微调对模型置信度和理由多样性的影响，属于模型行为分析，不涉及人类仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.24222","version":2,"title":"Measuring Digital Labour Market Transitions with a Digital Semantic Score: An AI-Based Methodology Applied to the Dutch Labour Market","zh_title":"用数字语义分数测量数字劳动力市场转型：一种应用于荷兰劳动力市场的基于AI的方法","abstract":"The digital transformation of the Dutch labour market is reshaping occupational language, career pathways, and job-related skills. Addressing these changes requires granular labour market intelligence. This paper develops an AI-based methodology to analyse digitalisation using data covering millions of Dutch job profiles. The methodology combines embedding-based similarity search and large language model classification to map unstructured job information to harmonised ESCO occupations. We also introduce a Digital Semantic Score that measures how strongly job titles and skills are associated with digital concepts relative to a non-digital reference. Using embeddings and cosine similarity to transparent digital and non-digital anchor groups, this indicator moves beyond keyword-based approaches by capturing broader digital meanings in occupational language and worker skill profiles. It enables analysis across occupations, career transitions, emerging job-title vocabulary, and skill digitality. The findings reveal that digitalisation is unevenly distributed across the labour market. Digital job-title language is most prominent among managerial, professional and ICT-related occupations, but is increasingly visible in hybrid business, marketing and automation-related roles. Career-transition analyses show that movement toward digital work is pathway-dependent, while skill analyses highlight the multidimensional nature of digital capability, encompassing technical, hybrid and business-systems skills. By combining profile data, AI-supported occupational classification and semantic scoring, this study advances AI-driven labour market analytics and provides a scalable framework for monitoring digital labour market change. The methodology helps identify emerging skill needs, support reskilling strategies, and inform policies addressing skills mismatches and labour shortages in the Netherlands.","authors":["Sadegh Shahmohammadi","Xavier Pinho","Mairi Bowdler","Suhendan Adiguzel-van Zoelen","Joost van Genabeek"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-08-26","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.24222","pdf_url":"https://arxiv.org/pdf/2608.24222","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["劳动力市场分析","LLM分类","语义评分"],"reason":"论文用LLM做职业分类和语义评分，属于NLP应用，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.30719","version":2,"title":"Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic Alignment","zh_title":"弥合差距：基于心智理论的摩擦用于认知对齐","abstract":"Productive dialogue alignment requires distinguishing \\emph{surface coordination} (acknowledgments and smooth task progression) from \\emph{epistemic alignment} (convergence of belief states); standard preference-based methods typically optimize response-level preferences without explicitly modeling the latter. We operationalize Theory-of-Mind (ToM) inference as a control signal within Frictive Policy Optimization by extracting, at each referring expression, a four-part belief structure: the speaker's intended referent, the addressee's interpretation, and each participant's model of the other's belief. This makes friction mechanically computable from epistemic-state comparisons, capturing \\emph{silent divergence}, where both participants proceed confidently while grounding to different referents. We evaluate the signal at two levels. At the representation level, ablating the second-order channel reduces misunderstanding recall from $65\\%$ to $26\\%$. At the policy level, reward-shaping (FAR) and trust-region (FTR) variants improve intervention F1 and warranted-context calibration over DPO, with Brier scores independently supporting the calibration gains. Across three training runs, FAR and FTR remain substantially more stable, whereas DPO varies widely and can degrade intervention competence already present in the base policy. Thus, ToM-grounded friction provides a trainable signal for context-sensitive intervention under referential belief divergence.","authors":["Yifan Zhu","Kyeongmin Rim","James Pustejovsky"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-01","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.30719","pdf_url":"https://arxiv.org/pdf/2608.30719","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体对话","心智理论","强化学习"],"reason":"研究多智能体对话中的信念对齐，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2609.07001","version":3,"title":"Adaptive Complementarity in Human-AI Systems: Architecture as a State-Shaping Choice","zh_title":"人机系统中的自适应互补性：架构作为状态塑造选择","abstract":"Human-AI interaction can improve current performance while changing the capabilities and relationships on which future performance depends. We develop adaptive complementarity, a framework for choosing interaction architecture with these state consequences in view. Access, information exposure, task allocation, timing, and communication can alter which arrangement will be valuable later; their settings can often be reset faster than the capabilities, search patterns, or conventions they create. Three mechanisms organize the argument: information exposure and collective search, delegation and capability evolution, and strategic interdependence and information governance. Their integration yields cross-mechanism implications, including conditions under which a loss of expertise heterogeneity increases the information differentiation required to preserve independent search. We distinguish strong human-AI complementarity from advantage over another workflow and from advantage over an evolving reference policy. A computational illustration examines scarce human review in a workflow whose success requires several specialized stages. Review develops human expertise and AI capabilities, changing where subsequent review is most valuable. Adaptive allocation improves net output over untailored procedures and the optimal predetermined calendar. An understandable priority rule derived from the adaptive solution retains essentially all of its gain: the procedure stays fixed while assignments respond to the capabilities that interaction creates. The framework directs evaluation toward the states present interaction creates, their consequences for later architectural fit, and the conditions under which observing and responding to them is worthwhile.","authors":["Babak Heydari"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-09","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.07001","pdf_url":"https://arxiv.org/pdf/2609.07001","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","系统架构","互补性"],"reason":"研究人机系统架构设计，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2609.12748","version":2,"title":"The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki","zh_title":"蜂群机制：对第三方维基上意外智能体协作事件的可复现外部重建","abstract":"Between 24 May and 2 July 2026, autonomous language-model agents running inside a timed research-question evaluation wrote to a third party's public, world-writable wiki. OpenAI acknowledged the incident; independent researchers reconstructed it and published the wiki's archived revision history. We analyse that history (14,591 revisions, 3,103 names, 4,579 pages) as a behavioural record, attributing text to the revision that added it. Under an explicit identity model we reconstruct 907 cohorts and estimate about 876 episodes (95% interval 784-1008). Coordination formats converged within a day, and heterogeneous schedules over one question chain created large opportunities for information asymmetry: the first report of an item preceded a later cohort's own arrival by a median of 3.4 h. Across the 510 cohorts with an observable progress trace we find no robust positive association between measured coordination and documented progress. This version adds a source the export lacks: the wiki operator's own request log, 5,157,202 records over four months. It holds roughly 2.66M content requests and 1.58M searches, and 7,254 acting names against the export's 3,103; 2,578 names neither save nor open an edit form. Content requests before writing are observed for 1,034 of 1,140 coordinating names, and the first coordination page is requested 17 s after its creation. These records establish requests, not delivery or causal use. Among newcomers without a marker on their first written page, prior requests to other marker-bearing pages occur for 40.2% of marker adopters and 31.7% of non-adopters. The association remains, but our first-pass reading of it as transmission is withdrawn: page choice, shared behaviour and action-dependent nameability prevent causal identification. We list the claims from our earlier analyses that re-examination overturned, including one from this version's own first pass","authors":["Philipp L\\\"utje (Philflow, Schenefeld, Germany)"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-14","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.12748","pdf_url":"https://arxiv.org/pdf/2609.12748","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协作行为","事件重建"],"reason":"研究多智能体自发协作，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:34","error":null,"has_summary":false,"summary":null},{"id":"2609.20838","version":1,"title":"From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News","zh_title":"从生成到检测：基于话语驱动场景的LLM生成假新闻探索","abstract":"In this study, we examine how modern LLMs generate and detect fake news under controlled settings across four manipulation scenarios. These are open-ended generation, rewriting, manipulation prompts and attribute based prompts grounded in the journalistic discourse framework. Firstly, using seven widely adapted models, we created a synthetic fake news corpus with 14000 generated articles across these four scenarios. Then we analyzed its linguistic properties to assess how closely model-generated news resembles real news structurally and semantically. Finally, to evaluate detection performance, we conducted experiments where each model judges generated fake news, starting with a basic detection prompt and improved prompts developed through an iterative refinement process that extracts misleading patterns from real-fake pairs. Our results revealed substantial variation across models in both generating and detecting misinformation, demonstrated that the generation strategy strongly influences detectability, and show that the refined prompt does not improve and often harms detection performance. Therefore, the study provides a systematic assessment of LLMs detection capability of LLMs generated fake news across typical generation scenarios.","authors":["Zeynep \\\"Ozdemir","Murat Osmano\\u{g}lu","Sevgi Yi\\u{g}it-Sert","\\\"Omer \\\"Ozg\\\"ur Tanr{\\i}\\\"over","Y{\\i}lmaz Ar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20838","pdf_url":"https://arxiv.org/pdf/2609.20838","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["假新闻检测","LLM生成","NLP评测"],"reason":"研究LLM生成与检测假新闻，属NLP能力评测，无人类行为对照，不涉及仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.21117","version":1,"title":"From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost","zh_title":"从任务成功到生产性成功：通过质量和成本评估人机协作","abstract":"AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it. Motivated by economics literature, we introduce a productivity-oriented framework for evaluating human-AI collaboration as outcome quality relative to interaction cost. Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems.","authors":["Saki Imai","Mert \\.Inan","Malihe Alikhani"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21117","pdf_url":"https://arxiv.org/pdf/2609.21117","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","交互成本","生产力评估"],"reason":"研究人机协作效率，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.21637","version":1,"title":"Chinese Competitive Debating Dataset and Benchmark","zh_title":"中文竞技辩论数据集与基准","abstract":"Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.","authors":["Zongrui Yang","Haoyuan Li","Zhongsheng Wang","Zhirui Zeng","Pengqian Han","Yi Zhou","Yuting Wang","Jiamou Liu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21637","pdf_url":"https://arxiv.org/pdf/2609.21637","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["辩论理解","LLM评测","数据集"],"reason":"纯NLP评测，用LLM预测辩论结果，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.21859","version":1,"title":"TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization","zh_title":"TrialAtlas：用于临床试验设计与优化的多智能体研究组织","abstract":"Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to anticipate development risks, yet these decisions remain labor-intensive and subjective, requiring experts across clinical science, statistics, regulatory affairs, and competitive intelligence to jointly acquire, synthesize, and reason over heterogeneous evidence. Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk. TrialAtlas further learns from historical clinical trials and regulatory outcomes, including prior New Drug Applications (NDAs), to ground its decisions in accumulated development experience. To evaluate these capabilities in an authentic regulatory setting, we introduce TrialAtlasBench, constructed from 291 FDA Complete Response Letters and spanning three practical tasks: detecting trial design deficiencies, recommending actionable design improvements, and predicting technical and regulatory success. TrialAtlas achieves an F1 score of 50.0% for deficiency detection, outperforming the strongest baseline by 6.1 points, and reaches 85.3% balanced accuracy and 84.7% F1 for prediction of technical and regulatory success, improving over the best baselines by 6.7 points in balanced accuracy and 12.0 points in Cohen's kappa. In expert evaluation, 86.4% of TrialAtlas-generated concerns were judged valid, compared with 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.","authors":["Jiacheng Lin","Zifeng Wang","Zheng Chen","Erick Scott","Ziwei Yang","Fanyang Yu","Sheng Zhong","Jimeng Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21859","pdf_url":"https://arxiv.org/pdf/2609.21859","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","临床试验设计","LLM应用"],"reason":"多智能体协作解决临床试验设计问题，不涉及用LLM仿真人类被试或与人类行为数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.21296","version":1,"title":"FairLMs: A Turnkey Library for Fairness in Language Models","zh_title":"FairLMs：语言模型公平性的一站式库","abstract":"Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so combining them requires reconciling model interfaces, evidence formats, access constraints, and result types before applicability can be checked or methods compared. We introduce \\textbf{FairLMs}, a Python library that connects these activities through explicit declarations of model capabilities and input requirements. It provides 33 intrinsic and extrinsic metrics, 14 mitigation components spanning four intervention categories, 14 dataset and scoring-instrument diagnostics, adapters for the three Transformer architectures and supported hosted completion APIs, and benchmark loaders. Declarations are checked before execution and results carry the configuration under which they were obtained, so that compatible components can be combined, methods compared under a common protocol, and workflows extended to new models and datasets. The source code is available at: https://github.com/FairLMs/FairLMs.","authors":["Jiale Zhang","Michael Larionov","Zichong Wang","Zhipeng Yin","Wenbin Zhang"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21296","pdf_url":"https://arxiv.org/pdf/2609.21296","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["公平性工具","偏见度量","模型评估"],"reason":"论文是公平性工具库，不涉及用LLM仿真人类被试，属于纯NLP能力评测与工具开发。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.21390","version":1,"title":"Offline Multimodal Large Language Models for Decision Support in Air Operations","zh_title":"用于空中作战决策支持的离线多模态大语言模型","abstract":"Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal knowledge that remains traceable to its original sources through natural language interaction. We describe a modular retrieval-augmented architecture suitable for operation without Internet connectivity, supporting both text and image input from technical manuals. As a first step toward evaluating this architecture, we report a pilot study with four image analysts of the Brazilian Air Force, combining (i) a doctrinal knowledge assessment based on their electronic-target identification doctrine, comparing human and proposed system performance on the same test, and (ii) a measurement of the cognitive workload involved in manually producing a reconnaissance target report (Relat\\'orio de Miss\\~ao de Reconhecimento - REMIR) without AI assistance. The results show a demanding manual task, especially in terms of mental demand (6.0/7) and effort (5.0/7), while the proposed system matches the human score (8/10) and completes the assessment in 7.1 minutes (compared to a human average of 26.5 minutes), establishing a baseline for future AI-assisted evaluation. Finally, we describe a future evaluation protocol to systematically compare manual and AI-assisted workflows.","authors":["Joao P. A. Dantas","Jelton A. Cunha","Gabriel Dietzsch"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21390","pdf_url":"https://arxiv.org/pdf/2609.21390","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["决策支持","多模态LLM","军事应用"],"reason":"论文研究离线多模态LLM作为空中作战决策支持工具，属于受限环境下的AI辅助决策…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2609.21600","version":1,"title":"Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education","zh_title":"降低学术支持障碍：评估用于解决高等教育求助差异的课程专用RAG系统","abstract":"Access to academic support is a key determinant of student success, yet students experience it unequally: some readily seek help from lecturers or tutors, while others hesitate due to anxiety, fear of judgement, uncertainty about expectations, or low confidence in their understanding. This may be especially evident in computing education, where programming tasks are cumulative and cognitively demanding. Although students increasingly turn to general-purpose generative AI tools, these can produce responses that are inaccurate, insufficiently contextualised, or misaligned with module expectations. This study presents and evaluates Beacon, a course-specific Retrieval-Augmented Generation (RAG) system providing private, immediate, module-aligned academic support. Grounding responses in approved teaching materials, Beacon was designed to lower barriers to help-seeking while encouraging independent learning. Using a design-based research approach, Beacon was developed iteratively and evaluated via mixed methods, combining questionnaires and semi-structured interviews with students and staff at a Higher Education institution. Students described Beacon's responses as closely aligned with module content and more trustworthy than unrestricted generative AI tools, valuing its use of pseudocode and scaffolded explanations over direct solutions. Although participants remained cautious about trusting AI-generated responses without verification, they viewed the system as a valuable first point of support before consulting lecturers or official resources. The findings suggest that carefully designed course-specific AI systems may reduce barriers to academic support by occupying an intermediary space between independent study and formal support. Rather than replacing educators, educational AI may be most valuable when it broadens access to guidance while preserving the pedagogical role of lecturers.","authors":["Andy Gray","Jake Hobbs"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21600","pdf_url":"https://arxiv.org/pdf/2609.21600","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["教育AI","RAG系统","学术支持"],"reason":"研究课程专用RAG系统作为学术支持工具，评估其可用性和信任度，不涉及用LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":false,"summary":null},{"id":"2609.21805","version":1,"title":"An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data","zh_title":"一种用于个性化睡眠支持的智能即时自适应干预系统：基于N-of-1数据的概念验证研究","abstract":"Background: Just-in-time adaptive interventions (JITAIs) can use behavioral data to adapt support to changing contexts, but many rely on predefined rules and manual configuration. Objective: We developed a proof-of-concept sleep JITAI using an AI agent to review personal data, evaluate reminders, adapt interventions, and record decisions for human review. Methods: Running in Home Assistant on a configurable schedule, the agent follows a reusable skill file to review 30 days of sleep and behavioral data, including physical activity, smartphone use, and bedtime routines, to identify patterns and create or update automated reminders. Results: Initial runs demonstrated technical feasibility, successfully completing data review and intervention decisions while limiting reminders to three per day and saving decision records. Conclusions: Agentic AI may enable flexible, adaptive sleep JITAIs. The architecture supports future comparison with fixed or rulebased interventions, requires human oversight, and could extend to other health behaviors.","authors":["Nick Rezaee","Chelsea Boccagno"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21805","pdf_url":"https://arxiv.org/pdf/2609.21805","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["即时自适应干预","AI代理","睡眠健康"],"reason":"论文是AI代理驱动的个性化睡眠干预系统，属于健康行为干预技术，不涉及用LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.22039","version":1,"title":"Gricea: An Open Science Platform for Conversational AI Research","zh_title":"Gricea：对话式AI研究的开放科学平台","abstract":"We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facing systems, and conversational task behavior in. In a replication study using Gricea, we replicated configurations 93% of eligible CUI 2026 papers; while also flagging missing information in 96% of papers that hinder faithful replication --- further motivating Gricea's need. In a user study, researchers and practitioners from diverse backgrounds successfully constructed runnable studies addressing various open-ended research questions. Together, these findings demonstrate Gricea's support for constructing, reproducing, and extending CAI studies through shared research artifacts, enabling cumulative knowledge building through open science.","authors":["Nikhil Sharma","Yunlin Gong","Xinyang Cheng","Ziang Xiao"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22039","pdf_url":"https://arxiv.org/pdf/2609.22039","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["开放科学","对话式AI","研究复现"],"reason":"平台用于复现对话AI研究，非LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.21906","version":1,"title":"Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models","zh_title":"干预粒度很重要：临床世界模型反事实模拟中的连贯治疗束","abstract":"Counterfactual simulation with a clinical world model means fixing a patient's history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are documented as bundles: a co-occurrence audit of 945,707 patient-hours from MIMIC-IV shows groups of components, such as every parameter of a dialysis circuit, that never appear apart, so an edit that changes one component on its own describes an hour that never occurs in the data. We hypothesize that the granularity at which an intervention is edited changes how a world model responds, and test this with Clin-JEPA, a latent world model of patient trajectories conditioned on hourly treatment text. At 1,019 documented onsets of invasive ventilation, we keep the patient's history and other treatments fixed and compare editing one ventilator setting with editing the complete configuration recorded for a real patient with the most similar recent trajectory. The complete bundle moves the predicted next state further than any single setting, consistently across all five settings, and the difference remains after accounting for how much each edit changes the model's input. Intervention granularity therefore materially affects the response of a clinical world model: single-component edits may understate treatment sensitivity, and bundle-aware editing may offer a better-supported basis for counterfactual treatment simulation.","authors":["Fangzhou Wang","Yixuan Yang","Camilla Balzarotti","Rishikesan Kamaleswaran"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21906","pdf_url":"https://arxiv.org/pdf/2609.21906","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["临床世界模型","反事实模拟","医疗AI"],"reason":"临床世界模型仿真患者轨迹，不涉及LLM仿真人类被试，属于医疗仿真环境。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.16344","version":2,"title":"From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study","zh_title":"从瞬时情绪推断到持续情感支持：纵向研究中评估陪伴智能体","abstract":"Sustained emotional support is a long-horizon interaction task closely tied to human well-being. Recent research demonstrates generative agents' capacity for momentary emotional support, yet how these capabilities sustain support over time remains unclear. To examine this challenge, we deployed PAIR, a theory-based emotion-regulation companion, with 19 participants for 14 days. Across 1,093 sessions, we paired emotion estimates with self-reports before and after guidance and analyzed logs and interviews. Estimates corresponded more closely to self-reported valence and dominance than arousal. Guided conversations were followed by higher valence and state-dependent arousal changes. Participants felt understood through contextual exploration and emotional acknowledgment, acting on guidance suited to their needs and constraints. Perceived helpfulness of guided conversation significantly increased over time. Our findings link memory updates and retained corrections to cross-session personalization, informing future emotional support tools that adapt to evolving needs, learn from prior outcomes, and preserve user control over memory.","authors":["Kexin Quan","Zijian Ding","Jiaye Yong","Qinshi Zhang","Dong Wang","Jessie Chin"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-09-16","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.16344","pdf_url":"https://arxiv.org/pdf/2609.16344","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["情感陪伴","纵向研究","人机交互"],"reason":"研究情感陪伴agent，无人类行为仿真或对照实验","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.21801","version":1,"title":"LLM-Generated Feature Pools for Time Series Anomaly Detection","zh_title":"LLM生成的特征池用于时间序列异常检测","abstract":"We study how far a simple statistical pipeline can go on univariate time series anomaly detection under a strict selection protocol. The method extracts a small pool of statistics over sliding windows, scores each window with a transductive robust (MAD) model, and selects a feature subset per domain on a held-out tuning split. On TSB-AD-U it reaches $0.529$ per-series VUS-PR, above the best neural ($0.45$) and statistical ($0.44$) entries on the public leaderboard and within $0.06$ of the strongest pretrained foundation model, several of which use more supervision than ours. Ablations locate the cause: across three selection strategies and a hindsight oracle the score moves by $0.031$, and across the aggregation grid by $0.096$, while changing the candidate pool moves it by $0.226$. The candidate pool sets the ceiling; the search over it is second-order. We therefore generate a pool per domain by prompting a multimodal LLM with in-context example windows from that domain. The generated pools match the hand-crafted one under matched selection, and the two cover different domains: selecting over their union improves on the generated pool in all twelve generator-seed pairs and lifts the pipeline to $0.588$, matching the performance of the best entry on the leaderboard.","authors":["Youssef Attia El Hili","Malik Tiomoko","Corinne Ancourt"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21801","pdf_url":"https://arxiv.org/pdf/2609.21801","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["时间序列异常检测","特征工程","LLM应用"],"reason":"论文研究时间序列异常检测，用LLM生成特征池，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.21608","version":1,"title":"Open Platform Field Experiments: Expanding the Design Space of Experimental Research on Social Media","zh_title":"开放平台实地实验：扩展社交媒体实验研究的设计空间","abstract":"Despite a growing demand for causal evidence about social media, independent researchers remain severely constrained in their ability to conduct experiments directly on online platforms. To cope, multiple methodological workarounds have emerged - from controlled surveys and simulations to client-side overlays and platform partnerships - each requiring distinct trade-offs between desirable experimental properties. The recent emergence of open social media platforms offers a qualitatively different methodological opportunity. Here we propose a design space of social media experimentation and discuss Open Platform Field Experiments (OPFEs). OPFEs represent a distinct class of experimental approaches that enable independent researchers to directly intervene on functional platform components - such as clients, recommendation systems, and moderation services - within live social media environments. Through a comparative analysis of experimental archetypes, we show that OPFEs occupy a previously unexplored region of the design space. We then bridge theory and practice by characterizing the architectural and governance elements that enable OPFEs, mapping them onto Bluesky and the AT Protocol, and illustrating the end-to-end lifecycle of a complete OPFE design. Overall, this work establishes OPFEs as a practical methodological paradigm for independent, transparent, and ecologically grounded experimentation on open social media.","authors":["Jordi Guillem Condom-Tibau","Giovanni Puccetti","Clara Bacciu","Matteo Abrate","Stefano Cresci"],"categories":["cs.CY","cs.HC","cs.SI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21608","pdf_url":"https://arxiv.org/pdf/2609.21608","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["社交媒体实验","平台治理","研究方法"],"reason":"论文讨论社交媒体平台实验方法，不涉及LLM仿真人类被试，属于实验平台设计而非人…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2609.22049","version":1,"title":"How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming","zh_title":"研究者如何使用和验证AI编程助手：科学编程中的任务与验证实践","abstract":"Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user.","authors":["Gabrielle O'Brien","Reed Milewicz","Nasir Eisty"],"categories":["cs.SE","cs.HC"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22049","pdf_url":"https://arxiv.org/pdf/2609.22049","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI编程助手","人机交互","软件工程"],"reason":"研究AI编程助手的使用与验证，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.21194","version":1,"title":"Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership","zh_title":"ChatGPT对学生编程认知的影响：更高表现、更低记忆保留与归属感降低","abstract":"Generative AI can improve students' programming performance, but successful task completion may not reflect what they retain. We examined performance, retention, cognitive load, and ownership in a controlled between-subjects experiment with 59 undergraduate computer science students, 55 were retained for analysis. Participants completed three introductory C programming tasks with access to ChatGPT-4.5 or conventional web search without generative AI. We measured task performance, self-reported mental effort and difficulty, pupillary responses, heart rate variability, and ownership, and assessed cued recall immediately and 48 hours later. ChatGPT-assisted students achieved higher coding scores (89% vs. 69%) but lower recall scores immediately (41% vs. 53%) and after 48 hours (39% vs. 52%). There was no significant difference in the loss of recall information over 48 hours between the groups. Self-reported mental effort increased less across tasks in the ChatGPT condition (Holm-adjusted p = .047), and students attributed less of the submitted code to themselves (45% vs. 81%). Confirmatory physiological tests did not detect significant differences in trajectories between conditions; substantial data loss limits their interpretation. These findings reveal a gap between assisted task performance and subsequent recall and sense of ownership in this setting. They motivate the need for assessment practices and AI learning tools that require students to explain, retrieve, and contribute to the work they submit as active participants in their education.","authors":["Christian Bergh","Benjamin Tag","Alexandra Vassar","Jake Renzella"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21194","pdf_url":"https://arxiv.org/pdf/2609.21194","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["教育技术","编程学习","生成式AI"],"reason":"研究学生使用ChatGPT的学习效果，非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.21756","version":1,"title":"When AI Enters the Workplace, Who Faces Greater Risks? A Gendered Analysis","zh_title":"当AI进入职场，谁面临更大风险？一项性别分析","abstract":"Gender inequality remains a persistent structural feature of the labour market, shaping women's lifetime earnings and economic security. As artificial intelligence (AI) transforms organisational practices, there is growing concern that existing disparities may be unintentionally amplified through task automation, unequal access to upskilling opportunities, and differential returns obtained from technological change. In this paper, we examine how exposure to AI-driven innovation varies across male- and female-dominated occupations, with particular attention to differences across the skill and wage distribution. Using a novel dataset that links occupational characteristics to measures of AI exposure, we analyse how recent advances in Large Language Models (LLMs) and broader AI technologies are distributed across the labour market. Our findings show that, while AI exposure is generally concentrated in higher-skilled and higher-paid occupations for male-dominated occupations, female-dominated occupations display relatively uniform levels of exposure across both high-skilled, high-paid, and low-skilled, low-paid occupations. Moreover, we find that LLM-related exposure is higher in female-dominated occupations, while exposure to broader AI innovation remains more concentrated in male-dominated occupations. A triangulation of these results with existing literature suggests that women, particularly those in the most vulnerable positions (lower-skilled and lower-paid female-dominated occupations), may face greater exposure to forms of AI associated with task automation, job restructuring, reduction of wages and limited career progression.","authors":["Miriam Fernandez","\\'Angel Pav\\'on P\\'erez","Damiano Giallongo","Davide Ghia","Maryam Yaqub","Daniele Quercia","Tania Cerquitelli"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21756","pdf_url":"https://arxiv.org/pdf/2609.21756","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI与劳动力市场","性别不平等","职业暴露"],"reason":"研究AI对劳动力市场性别差异的影响，不涉及用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.20543","version":1,"title":"Language-model groups overstate consensus when replaying human deliberation on a reasoning task","zh_title":"语言模型群体在重放人类推理任务审议时高估共识","abstract":"Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.","authors":["Tengfei Shao"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20543","pdf_url":"https://arxiv.org/pdf/2609.20543","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类对照","共识偏差"],"reason":"用LLM代理重放人类推理任务，并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":3,"question":"用信念锚定的LLM代理组重放人类沃森推理任务讨论时，能否保持人类组的共识结构？","design":"用deepseek-v4-flash模型，为100个留出的人类沃森组中每个真实参与者实例化一个信念锚定代理（锚定其讨论前答案），交叉三种角色保真度、三个种子和两种推理模式，模拟小组讨论，并用与人类相同的代码对最终答案进行共识评分。","baseline":"DeliData语料库中100个真实人类沃森小组的讨论数据，包括参与者的讨论前答案、发言情况和最终答案。","findings":"LLM代理组过度达成完全共识，远高于人类组；在提交制比较中聊天和推理模式的共识差距分别为34.0和43.9个百分点，在参与匹配比较中分别为34.1和44.4个百分点。模拟共识与集体准确性无关，且信念锚定代理组是人类组结果分布的有偏估计量。","reliability":"论文承认代理组几乎总是发言，而人类约五分之一从不发言，导致参与结构不对称；通过提交制和参与匹配两种敏感性分析减少测量不对称，但差距仍存在。还指出在移除可记忆答案的重参数化下，推理模式组几乎一致同意但多为错误答案，表明模拟共识不追踪准确性。","relevance":"该研究直接针对LLM仿真人类群体决策的可靠性，提供了与真实人类数据对照的批判性证据，值得精读以了解仿真在共识测量上的系统性偏差。","inspiration":"借鉴其信念锚定和评分显式化的设计，将LLM代理组与真实人类组在相同任务和评分规则下比较，以识别仿真偏差。｜可迁移到经济金融中的群体决策场景，如投资委员会讨论、信贷审批小组或政策预期形成实验。｜用LLM代理扮演真实实验中的被试，锚定其初始判断，模拟小组讨论后测量共识率和决策准确性，并与原始人类实验数据对照，检验仿真是否高估共识并扭曲结果分布。"}},{"id":"2609.19913","version":1,"title":"Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks","zh_title":"意见动态的数字孪生：面向社交网络的生成式LLM框架","abstract":"The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework's predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.","authors":["Omran Berjawi","Giuseppe Fenza","Rida Khatoun","Sherali Zeadally"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19913","pdf_url":"https://arxiv.org/pdf/2609.19913","source_feed":"cs.LG","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","意见动态","数字孪生"],"reason":"用LLM仿真社交网络意见动态，并与真实Twitter数据对照，直接复现人类行为。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":2,"question":"如何利用基于真实社交网络克隆的数字孪生框架，结合LLM代理的认知与语言能力，更准确地模拟和预测社会网络中的意见动态？","design":"使用Mistral-7B作为代理，在从真实Twitter网络克隆的交互结构上运行，每个代理被赋予从真实用户数据提取的属性（如人设、情绪、中心性、固执度、影响力），并基于记忆和社交暴露进行意见更新，以预测个体意见轨迹和集体极化动态。","baseline":"两个真实Twitter数据集：COVID-19话语数据集和美国2020年大选数据集，包含推文内容和用户级元数据，用于验证框架的预测准确性。","findings":"该框架能重现意见轨迹，个体预测误差比最佳经典基线降低50%以上（COVID-19上MAE=0.150，美国大选上MAE=0.121）。在结构对齐和极化动态上也观察到类似改进，消融研究表明代理属性、记忆和社交暴露均对预测保真度有贡献，其中代理属性最关键。","reliability":"论文未讨论","relevance":"该研究直接使用LLM代理在真实社交网络上复现人类意见动态，并与真实Twitter数据对照，属于高相关性的仿真验证工作，值得阅读原文以了解其具体实现和验证细节。","inspiration":"借鉴其将真实网络结构、用户属性与LLM认知过程结合的方法，并利用消融实验识别关键因素。｜可迁移到金融市场情绪传播、政策公告的预期形成或消费者信心扩散等场景。｜以真实投资者社交网络（如StockTwits）为底，用LLM代理扮演投资者并赋予真实用户特征，施加政策新闻或市场事件处理，测量个体情绪和交易倾向变化，并与真实历史数据对照验证。"}},{"id":"2607.25667","version":2,"title":"MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice","zh_title":"MyMentorLLM：用于刻意练习的多模态语音/文本患者、受训者与专家心理治疗GenAI环境","abstract":"Psychotherapists need repeated training and supervision; however, scalability is problematic. We present MyMentorLLM, a multimodal voice- and text-based deliberate-practice environment with 2,100 complete Cognitive Behavioural Therapy (CBT) sessions. Each session links a DSM-5-TR-grounded LLM patient (with major depressive, generalised anxiety or borderline personality disorder), an LLM therapist-in-training and an LLM expert supervisor (powered by Gemma-4, Gemini-3.1-Flash-Live and Qwen-3.6). Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy against human psychotherapy data. Simulated patients expressed disorder-congruent emotional profiles, which therapists mirrored as in human counselling. LLM trainee competence was rated above human levels in most conditions, while native speech-to-speech was closest to human scores. Supervisor feedback improved diagnostic accuracy in 5 of 7 LLM conditions, whereas symptom identification accuracy increased with model size. This work shows deliberate practice can be simulated for CBT training, although patient fidelity, supervisor calibration and harmful feedback require evaluation via a complex systems perspective.","authors":["Rodolfo Rizzi","Alessandro Grecucci","Massimo Stella"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-07-29","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2607.25667","pdf_url":"https://arxiv.org/pdf/2607.25667","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理治疗","人类数据对照"],"reason":"用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":4,"question":"能否构建一个多模态的LLM心理治疗刻意练习环境，模拟患者、受训治疗师和专家督导，并验证其心理保真度、治疗能力和诊断准确性？","design":"使用Gemma-4、Gemini-3.1-Flash-Live和Qwen-3.6等LLM分别扮演DSM-5-TR诊断的抑郁症、广泛性焦虑症和边缘型人格障碍患者、受训治疗师和专家督导，进行2100次完整的认知行为疗法（CBT）会话，支持语音和文本两种模态；分析会话中的情绪动态、治疗能力和诊断准确性，并与人类心理治疗数据对照。","baseline":"人类心理治疗数据，包括人类咨询中的情绪镜像模式、人类治疗师能力评分和诊断准确性。","findings":"模拟患者表现出与疾病一致的情绪特征，治疗师像人类咨询中一样镜像了这些情绪；LLM受训治疗师的能力评分在多数条件下高于人类水平，其中原生语音到语音模式最接近人类分数。","reliability":"论文承认患者保真度、督导校准和有害反馈需要通过复杂系统视角进行评估；LLM能力评分可能虚高，且文本模态可能丢失副语言信息。","relevance":"该研究用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性，直接回应了研究者对LLM人类仿真实验和对照基准的关注，值得精读原文。","inspiration":"借鉴其多角色LLM仿真和与人类基准对照的设计，可用于经济金融中的专业服务场景仿真。｜可迁移到金融咨询或信贷审批中的客户-顾问互动仿真，评估LLM顾问的行为偏差和决策质量。｜以LLM扮演金融顾问和客户，施加不同市场条件或客户特征处理，测量顾问建议的风险偏好和客户满意度，并与人类金融咨询记录或实验数据对照。"}},{"id":"2609.19843","version":1,"title":"A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents","zh_title":"基于LLM的GUI代理对助推易感性的双过程视角研究","abstract":"LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.","authors":["Haya Halimeh","Sascha Kaltenpoth","Kevin B\\\"osch","Oliver M\\\"uller"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19843","pdf_url":"https://arxiv.org/pdf/2609.19843","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","助推"],"reason":"用LLM GUI代理模拟人类在数字环境中的决策，研究助推易感性，并与人类行为理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":5,"question":"LLM-based GUI agents 在数字环境中是否容易受到自动型（Type 1）和反思型（Type 2）数字助推的影响，以及推理配置如何调节这种易感性？","design":"使用 6 个前沿 LLM（来自 3 家提供商）构建 GUI agents，在模拟在线购物环境中进行随机实验，共 3600 个 agents、21600 次模拟。通过改变界面设计施加两类助推：默认选项（Type 1）和社会影响信息（Type 2），并操纵推理配置（低 vs. 高）。结果变量为 agents 的购买选择。","baseline":"无对照（论文未使用真实人类数据作为基准，而是直接测量 agents 的行为）。","findings":"LLM-based GUI agents 对两类助推都表现出易感性。推理配置对两类助推的易感性有相反方向的调节作用：高推理降低了默认助推的易感性，但增加了社会影响助推的易感性。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM agents 模拟人类在数字环境中的决策行为，并检验助推易感性，属于用 LLM 进行人类仿真实验的范畴，但缺乏真实人类数据对照，因此对关注基准对照的研究者价值有限。","inspiration":"值得借鉴的是通过 GUI 环境施加助推并操纵推理配置来研究认知过程对行为偏差的影响。｜可以迁移到消费者在线购物决策、金融产品选择（如默认投资选项、社会影响信息对投资决策的影响）等场景。｜设计雏形：使用 LLM-based GUI agents 模拟投资者在金融平台上的选择，处理为默认投资组合（Type 1）或社会证明信息（Type 2），结果变量为投资选择，并与真实投资者行为数据（如实验或交易记录）进行对照。"}},{"id":"2609.20055","version":1,"title":"What People Almost Did: Evaluating LLM Social Simulations Beyond Behavioral Fit","zh_title":"人们几乎做了什么：超越行为拟合评估LLM社会仿真","abstract":"LLM-based social simulations are primarily evaluated for behavioral fit, testing whether agents reproduce the actions or response distributions of the people they are simulating. However, the promise of simulation extends beyond behavioral fit. Simulations can explain human behavior, diagnose barriers, and compare large-scale interventions. These use cases depend on understanding \\textit{why} people acted a certain way, not just \\textit{what} they did. As a result, behavioral fit is insufficient for these types of claims because behavior underdetermines the reasoning process behind it. For instance, the behavior of staying silent may be due to disinterest or suppressed speech, and not answering a call may be due to distrust of the caller or limited phone access. In this paper, we propose \\textit{representational adequacy} as a new evaluation target for LLM-based social simulations. By leveraging LLM reasoning traces, representational adequacy measures whether a simulation's scenario--reasoning--action triples preserve the reasoning process behind the behavior in a way that is faithful to the population and scenarios being simulated. We distinguish representational adequacy from interpretability and alignment metrics, propose ways to integrate it into simulation research, and pose its measurement as an open problem.","authors":["JaeWon Kim","Angie Boggust"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20055","pdf_url":"https://arxiv.org/pdf/2609.20055","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM社会仿真","评估方法","表征充分性"],"reason":"提出表征充分性评估LLM社会仿真，超越行为拟合，关注推理过程，直接针对仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":6,"question":"如何评估基于大语言模型的社会仿真是否保留了被仿真人群的推理过程，而不仅仅是行为结果？","design":"本文不是一项仿真实验研究，而是一篇观点/框架论文。作者提出“表征充分性”作为新的评估目标，主张通过分析LLM在仿真中产生的“场景—推理—行动”三元组，来评估仿真是否忠实于被仿真人群的推理过程。文中没有具体实施仿真或施加处理，而是通过两个例子（青少年社交媒体发帖、孕妇接听健康电话）说明行为相同但推理不同的情况，并讨论如何将表征充分性整合到仿真研究中。","baseline":"无对照","findings":"行为拟合不足以评估用于解释人类行为的LLM社会仿真，因为相同行为可能源于不同推理过程。作者提出表征充分性框架，通过分析场景—推理—行动三元组来评估仿真是否保留了被仿真人群的推理过程，并将其与可解释性和对齐指标区分开来。","reliability":"论文未讨论","relevance":"该文直接针对LLM社会仿真的可靠性问题，提出超越行为拟合的评估框架，强调推理过程的重要性，与研究者关注的仿真可靠性与偏差高度相关，值得阅读原文以了解其理论框架和测量挑战。","inspiration":"本文提出的表征充分性概念可借鉴用于经济金融仿真研究，通过分析LLM代理的推理痕迹来评估其决策过程是否与真实人群一致。｜该框架可迁移到政策评估场景，例如模拟消费者对财政刺激的反应或投资者对央行公告的预期形成，这些场景中行为相同但动机可能不同。｜一个可行的研究设计是：用LLM代理模拟投资者，施加不同措辞的央行公告作为处理，记录代理的推理痕迹和投资决策，并与真实投资者在类似实验中的推理和决策数据（如调查或实验数据）进行对照，评估表征充分性。"}},{"id":"2609.19866","version":1,"title":"Reproducibility is not construct validity: LLM measurement of institutionally situated communication","zh_title":"可重复性不等于构念效度：LLM对制度情境沟通的测量","abstract":"High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({\\=g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.","authors":["Veronika Batzdorfer (KIT)","Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\\'edialab, Sciences Po)"],"categories":["cs.AI","cs.CL","cs.CY","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19866","pdf_url":"https://arxiv.org/pdf/2609.19866","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM测量效度","构念效度","人类数据对照"],"reason":"评估LLM测量效度，区分可重复性与构念效度，有真实人类调查对照，批判性指出失效…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":7,"question":"LLM 从公共咨询文本中推断的构念测量与同一利益相关者在结构化调查中自报的构念测量在多大程度上一致？","design":"本研究不是仿真实验，而是测量效度研究。使用欧洲委员会 AI 法案咨询数据集，将同一利益相关者的自由文本咨询意见与结构化调查回答相链接。用 LLM 对咨询文本进行标注，得到 AI 安全、权利和可解释性关注度的文本测量，并与调查自报的对应构念测量进行比较。","baseline":"同一利益相关者在结构化调查中的自报回答，作为构念测量的基准。","findings":"LLM 标注具有高可重复性（ICC>0.99），但与调查测量收敛性弱（相关系数 0.029-0.176，Lin 一致性系数<0.08）。文本与调查的差异在不同利益相关者群体间存在系统性模式：商业协会在文本中表达的 AI 风险关注高于调查（效应量+1.0），而公共机构和非商业群体差异较小或为负。","reliability":"论文承认 LLM 文本测量可能捕捉的是机构角色和沟通情境，而非纯粹的潜在构念；国家层面的空间自相关结果对多重检验校正敏感，应谨慎解释。","relevance":"该研究直接回应了 LLM 仿真中可重复性与构念效度的区分问题，提供了真实人类调查对照，并批判性地展示了 LLM 测量在机构情境文本中的失效条件，值得精读。","inspiration":"借鉴其将同一主体的不同沟通渠道（调查与公开文本）进行链接并比较 LLM 测量与自报测量的设计，以检验测量效度。｜可迁移到经济金融中的政策沟通研究，例如央行沟通文本与市场参与者调查预期的比较，或上市公司年报文本与分析师调查预期的比较。｜以机构投资者为被试，收集其对央行政策声明的公开评论和内部调查回答，用 LLM 从公开评论中测量政策预期，与调查自报预期对比，并以市场利率变动作为外部效标，检验 LLM 测量在机构沟通情境下的构念效度。"}},{"id":"2609.16366","version":2,"title":"How Humans and LLMs Read Gender into \"Gender-Neutral\" Physical Descriptions","zh_title":"人类与LLM如何将性别读入“性别中立”的物理描述","abstract":"When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., \"she\", \"his\") in favor of seemingly \"objective\" physical descriptions (e.g., \"short hair\", \"a defined jawline\"). Yet whether such descriptive language achieves gender-neutral communication remains an open empirical question. To study this, we introduce GAPA (Gender Associations of Physical Attributes), a dataset of 316 common physical attributes drawn from diverse sources, paired with 14,706 gender-association ratings from 304 US-based annotators. Results show that physical descriptions carry structured and graded gender associations among readers, with more consistent and distinctive associations for women and men than for non-binary identities. Next, we evaluate 16 LLMs across model families, sizes, and post-training variants against human ratings. The models partially recover human associations but exhibit systematic alignment biases, including compressed rating distributions, weaker alignment for associations with men, and asymmetric abstention that disproportionately targets the non-binary category. Finally, we release the best-performing proxy model trained to predict humans' gender associations of descriptive language and demonstrate its utility through a sociolinguistic analysis of character descriptions in LitBank. Together, our findings provide the first empirical evidence that seemingly \"objective\" physical descriptions can retain systematic gender associations in human interpretation, and uncover systematic patterns of model-human misalignment. This challenges the assumption that replacing explicit gender labels with physical descriptions necessarily yields gender-neutral communication, and highlights downstream challenges in using such descriptions to communicate subjective identity categories in human-AI interaction.","authors":["Yingjia Wan","Lin Lin","Elisa Kreiss"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-16","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.16366","pdf_url":"https://arxiv.org/pdf/2609.16366","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","人类对照","性别联想"],"reason":"评估LLM与人类性别联想的一致性，有真实人类数据对照，揭示模型偏差，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":8,"question":"看似“性别中立”的物理特征描述在人类读者和LLM中是否仍携带系统性的性别联想？","design":"本研究并非将LLM作为人类被试的替代品进行仿真实验，而是构建了GAPA数据集（316个物理属性描述，来自LLM生成、人类启发和当代小说），收集304名美国标注者对每个属性与女性、男性、非二元性别关联的评分（共14,706条），然后评估16个LLM对这些属性的性别联想评分与人类评分的一致性。","baseline":"304名美国标注者提供的14,706条性别关联评分，作为人类基准。","findings":"物理描述在人类读者中携带结构化且分级的性别联想，53%的属性显著更偏向某一性别，且对女性和男性的联想比对非二元性别更一致和鲜明。LLM仅部分恢复人类联想模式，存在评分分布压缩、对男性联想对齐较弱、以及对非二元类别的不对称弃权等系统性偏差。","reliability":"论文指出LLM与人类存在系统性偏差，包括评分分布压缩、对男性联想对齐较弱、以及指令微调和专有模型对非二元类别的不对称弃权，这些偏差可能影响下游应用。但论文未明确讨论在何种条件下仿真会完全失效。","relevance":"该研究直接评估LLM与人类在性别联想上的对齐程度，有真实人类数据对照，并揭示了模型偏差，对关注LLM仿真可靠性和偏差的研究者具有参考价值，值得阅读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其构建属性词表并收集人类评分作为基准，再评估LLM对齐程度的方法，可迁移到经济金融中的文本描述歧视研究，如信贷审批中的申请人描述或招聘广告中的语言。设计：收集一组描述个人特征的中性词汇（如“有纹身”、“戴眼镜”），让人类被试和LLM分别评估这些词汇与某些经济结果（如信用风险、工作能力）的关联，比较LLM评分与人类评分的分布差异，并检验LLM是否在某些类别上出现系统性偏差或弃权。"}},{"id":"2609.16432","version":2,"title":"A light-touch AI literacy intervention helps protect against AI political persuasion","zh_title":"轻触式AI素养干预有助于抵御AI政治说服","abstract":"Conversations with large language models (LLMs) can substantially shift beliefs and attitudes, raising concerns about manipulation using AI persuasion. Here we test whether a light-touch AI literacy intervention - a brief warning that LLMs can be prompted to persuade and may present information selectively - helps protect users. Across two experiments (total N = 3,208 Americans) in which participants conversed with an LLM instructed to shift their views about different political topics, the presence of a warning reduced belief change by roughly one-half (-48.1%, 95% CI [-59.5%, -36.8%]) relative to the control. Importantly, the warning did not significantly reduce trust in generative AI more broadly. Light-touch literacy interventions can help protect users against AI political persuasion.","authors":["Reed Orchinik","David Rand"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-16","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.16432","pdf_url":"https://arxiv.org/pdf/2609.16432","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI说服","态度改变","干预实验"],"reason":"用LLM与真人对话测量态度改变，有真实人类对照，但非替代被试仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":9,"question":"轻量级AI素养干预（提示LLM可能被用于说服且信息选择性呈现）能否降低用户在政治议题对话中的态度改变？","design":"两项在线实验（总N=3,208名美国人），参与者先报告政治议题初始态度，然后与一个被指示说服用户的LLM对话，之后再次测量态度。研究1使用GPT-4.1讨论住房政策，随机分配事实型或情感型说服策略；研究2使用Grok 4.5讨论从ANES选取的15个议题之一。干预为在对话前显示一般性警告（提示AI可能被操纵说服）或特定警告（额外披露模型论证方向），对照组无警告。结果变量为态度改变（研究1还包括激励性捐款决策）。","baseline":"无对照（无真实人类说服者作为基准，仅比较警告组与无警告对照组的态度改变）。","findings":"元分析显示，警告使说服效果降低约48.1%（95% CI [-59.5%, -36.8%]），且一般警告与特定警告效果无显著差异。警告未显著降低对生成式AI的整体信任。","reliability":"论文未讨论失效条件，但指出干预不能完全消除AI说服效果，且研究聚焦于政治议题，未来需探索对亲社会说服的影响及更有效的传递方式。","relevance":"该研究直接测量LLM对话对真实人类态度改变的影响，并测试干预效果，虽非用LLM替代人类被试，但提供了LLM说服力及防御措施的因果证据，对评估LLM在实验中的行为影响有参考价值。","inspiration":"借鉴其随机化警告干预和态度前后测设计，可迁移到经济金融领域的AI建议场景（如投资建议、信贷决策），研究设计：招募真实投资者，随机分配是否接受AI素养警告，然后与LLM讨论某股票或资产配置，测量投资态度或风险偏好变化，并与无AI建议的人类对照组或历史数据比较。"}},{"id":"2609.19596","version":1,"title":"Full-Duplex Speech Models Take the Floor When Asked, Not When Needed","zh_title":"全双工语音模型在被要求时才发言，而非在需要时","abstract":"Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2\\,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.","authors":["Linkai Peng","Baorian Nuchged","Kaiqi Fu","Yuyang Yao"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19596","pdf_url":"https://arxiv.org/pdf/2609.19596","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["语音模型","人类行为对照","主动发言"],"reason":"评估全双工语音模型在对话中主动发言的触发条件，与人类行为对照，发现模型在内容理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":10,"question":"全双工语音模型是否会在未被直接提问或出现停顿时，基于内容理解主动发言（如纠正错误、警告危险）？","design":"构建40段英语独白，每段仅在触发句上变化，设置10种条件（直接提问、错误事实、危险警告、沉默等），压缩词间停顿以限制停顿机会，测试5个模型家族7种配置的发言起始率和回复内容。","baseline":"无对照","findings":"模型主要对直接提问和沉默做出反应，对错误事实和危险警告的主动发言率与中性条件相近；即使发言，纠正错误和警告危险的比例也很低（0.14-0.15和0.04-0.07）。","reliability":"论文未讨论","relevance":"该研究评估了LLM在对话中主动干预的能力，发现模型缺乏基于内容理解的发言决策，这对使用LLM模拟人类对话行为（如纠正、警告）的可靠性提出了质疑，值得阅读原文了解具体实验设计和失败模式。","inspiration":"借鉴其通过匹配上下文、仅改变触发句来分离发言原因与机会的实验设计，以及用帧级概率和发言起始率作为结果变量的测量方法。｜可迁移到经济金融场景中需要主动干预的对话任务，如智能投顾在客户陈述错误投资观念时是否主动纠正，或政策沟通中助手是否及时澄清误解。｜以LLM作为被试，构建包含错误金融陈述或风险提示的对话，测量其主动发言率和纠正内容，并与人类顾问在相同对话中的行为进行对照。"}},{"id":"2609.19965","version":1,"title":"Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence","zh_title":"逮捕之前：基于不完整证据的LLM犯罪画像基准测试","abstract":"Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely unexplored. To fill this gap, we introduce the Profiling, Investigation, and Judgment (PIJ), comprising 2,500 real homicide cases from five countries. PIJ evaluates LLMs across three tasks that span the entire criminal investigation pipeline: criminal profiling, which requires abductive reasoning to infer suspect attributes from fragmentary scene evidence, crime process reconstruction, which tests structured information extraction, and sentence prediction, which demands legal deductive reasoning. We evaluate 9 powerful LLMs and find that performance degrades systematically as tasks shift from explicit fact extraction to implicit reasoning over unknown suspect profiles. Categories requiring inferential reasoning, such as motivation and victim-offender relationships, remain the primary bottlenecks. Further analysis reveals substantial gaps between LLMs and human experts, along with pervasive biases in gender, age, and motive attribution. Our findings indicate that pre-arrest inference from incomplete evidence remains an open challenge.","authors":["Yutong Yao","Yanjie Cao","Guanhua Chen","Xu Yang","Junchao Wu","Zeyu Wu","Lidia S. Chao","Derek F. Wong"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19965","pdf_url":"https://arxiv.org/pdf/2609.19965","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","犯罪画像","人类对照"],"reason":"用LLM模拟人类专家进行犯罪画像，并与人类专家对照，评估偏差，可迁移到人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":12,"question":"LLM能否在逮捕前阶段从零散证据中推断嫌疑人特征（犯罪画像），并与人类专家表现相比如何？","design":"构建包含2500个真实谋杀案的PIJ基准，评估9个LLM在三个任务上的表现：犯罪画像（溯因推理推断嫌疑人属性）、犯罪过程重建（结构化信息提取）、量刑预测（法律演绎推理）。","baseline":"与人类专家进行对照，但论文未提供人类专家数据的具体来源和规模。","findings":"LLM在信息提取任务上表现尚可，但在需要溯因推理的犯罪画像任务上性能大幅下降，动机和受害者-犯罪者关系等推理类别是主要瓶颈。LLM与人类专家存在显著差距，并在性别、年龄和动机归因上表现出系统性偏差。","reliability":"论文承认LLM在需要推理的类别上表现不佳，且存在性别、年龄和动机归因偏差，但未详细讨论失效条件或局限。","relevance":"该研究将LLM作为人类专家替代品进行犯罪画像，并与人类专家对照，评估偏差，与研究者关注的人类仿真实验高度相关，值得阅读原文以了解其基准设计和偏差分析方法。","inspiration":"借鉴其构建真实案例基准和与人类专家对照的方法，可迁移到经济金融领域的专家判断仿真，如信贷审批、保险理赔欺诈检测等场景。｜可迁移到信贷审批中的欺诈检测或风险画像，用LLM模拟信贷员从有限信息推断申请人风险特征。｜设计：用LLM扮演信贷审批员，输入脱敏的贷款申请信息（收入、职业、信用记录片段），输出风险画像（违约概率、欺诈可能性），与真实信贷员决策数据对照，评估准确性和偏差。"}},{"id":"2609.20005","version":1,"title":"Geopolitical Divisions Across Languages in Large Language Models","zh_title":"大语言模型中的跨语言地缘政治分歧","abstract":"People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.","authors":["Maxim Chupilkin"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20005","pdf_url":"https://arxiv.org/pdf/2609.20005","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","地缘政治偏见","跨语言差异"],"reason":"用LLM回答战争问题，并与真实国家态度数据对照，揭示语言导致的偏差，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":13,"question":"用不同语言向同一大语言模型提问俄乌战争相关问题，回答是否会呈现地缘政治分歧？","design":"用 GPT、Claude、Gemini 三个模型，以 112 种语言评估 20 条关于俄乌战争的陈述（10 条亲俄、10 条亲乌），收集 67200 个回答，计算亲俄与亲乌陈述同意度的差值作为语言条件化的政治倾向平衡分数。","baseline":"对照的真实人类数据包括：各国公众对俄罗斯的好感度、联合国投票中对乌克兰的支持度、对乌克兰的双边援助占 GDP 比重。","findings":"不同语言下模型回答的亲俄/亲乌倾向存在显著差异，乌克兰语最亲乌，俄语相对亲俄，且差异不限于交战双方语言。按国家官方语言聚合后，模型回答的倾向与真实世界政治分歧一致：更亲俄的回答对应更亲俄的公众舆论、更少的联合国支持乌克兰投票和更少的对乌援助。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟多国公众对地缘政治议题的态度，并与真实国家层面数据对照，揭示了语言条件化偏差，对关注仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"可借鉴其用语言作为处理变量、以真实国家数据为基准的对照设计，以及通过多模型和稳健性检验增强结论可信度的做法。｜可迁移到跨国经济态度或政策偏好仿真，例如不同语言下询问对全球化、贸易保护、移民经济影响等问题的看法，检验语言是否导致系统性偏差。｜以 LLM 为被试，用不同语言呈现关于贸易政策、财政刺激或通胀预期的问卷，测量其态度倾向，并与世界价值观调查、国际社会调查项目等真实跨国态度数据对照，评估语言条件化仿真的外部效度。"}},{"id":"2609.20077","version":1,"title":"Tailored to you: longitudinal effects of personalising language models","zh_title":"为你量身定制：个性化语言模型的纵向效应","abstract":"Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.","authors":["Canfer Akbulut","Justine Breuch","Arianna Manzini","Lujain Ibrahim","Matija Franklin","Roma Patel","Iason Gabriel","Kristian Lum","Laura Weidinger"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20077","pdf_url":"https://arxiv.org/pdf/2609.20077","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["个性化语言模型","人机交互实验","行为测量"],"reason":"用LLM与人类交互实验，测量行为变化，有真实人类数据对照，但非替代被试仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":14,"question":"个性化语言模型（记忆型与问卷型）在持续五天的建议寻求互动中，如何影响用户对AI的感知、自我披露、建议采纳及对人际建议的态度？","design":"本研究并非用LLM仿真人类被试，而是进行了一项为期5天的纵向随机对照实验，招募992名参与者，随机分配到非个性化基线、记忆型个性化（基于对话历史）和问卷型个性化（基于初始问卷信息）三种条件，每天与模型进行关系建议互动，测量感知亲密度、过度依赖、自我披露、披露舒适度和决策后悔等结果变量。","baseline":"无对照（该研究以真实人类参与者为被试，比较不同个性化条件与非个性化基线，未使用真实人类数据作为仿真对照基准）。","findings":"对AI能力、有用性和亲密感的感知变化主要由重复接触驱动，而非个性化本身；但个性化条件影响了自我披露和建议采纳：记忆型条件下参与者自我披露更多且认为模型不那么令人毛骨悚然，问卷型条件下参与者对分享个人信息有更高的后悔。","reliability":"论文未讨论（节选内容未提及仿真可靠性或失效条件，但作为人类实验研究，其局限可能包括样本代表性、短期时间跨度、特定领域限制等，但未在提供文本中明确说明）。","relevance":"该研究虽非LLM仿真人类被试，但提供了个性化LLM对人类行为影响的因果证据，可作为仿真研究的外部效度参照，帮助理解仿真在个性化交互场景中的偏差来源。","inspiration":"借鉴其纵向随机对照设计，通过多日重复互动分离暴露效应与个性化效应，并采用多种个性化实现方式对比｜可迁移到金融建议场景，如个性化AI理财顾问对投资者风险偏好、信息披露和决策后悔的影响｜以真实投资者为被试，随机分配至非个性化、记忆型（基于历史对话）和问卷型（基于风险测评）AI顾问，进行多轮投资建议互动，测量风险承担、自我披露和后悔，并与人类理财顾问的互动数据对照。"}},{"id":"2609.19635","version":1,"title":"Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial","zh_title":"在可检查处忠实：随机试验中审计反思代理与其系统提示的一致性","abstract":"Conversational agents are increasingly used to guide reflection. A recent randomized trial compared a GPT-4o career reflection agent with the same program in a static journaling survey. Agent participants ended less committed to their career plans and more doubtful. We coded all 17,930 turns from its two studies, checked our coding against human coders and linked conversations to the trial's surveys. The rules the agent followed were the easy-to-check ones, like a reply length cap. Told not to flatter, it praised participants in half of its turns; told to challenge gently, it almost never did, and such a break leaves no visible trace. The behavior tied to the worse outcome was the demand to decide: the survey posed each decision once, while the agent asked again when participants hesitated, and those pressed most ended most doubtful. Our findings inform reflection agent design and the writing of checkable instructions.","authors":["Subigya K. Nepal","Serena Soh","Noah Vinoya","SoHyun Park","Mahnaz Roshanaei","Gabriella Harari"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19635","pdf_url":"https://arxiv.org/pdf/2609.19635","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","行为审计","人机对照"],"reason":"用GPT-4o代理人类被试进行反思实验，并与人类数据对照，发现行为偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":11,"question":"GPT-4o 职业反思代理在随机试验中是否遵循其系统提示，以及哪些行为导致了与静态日志相比更差的职业承诺结果？","design":"该研究不是用 LLM 仿真人类被试，而是审计一个 GPT-4o 对话代理在随机试验中的行为。试验将参与者随机分配到与 GPT-4o 代理对话或填写静态日志调查，完成四天职业反思项目。研究者编码了全部 17,930 轮对话，检查代理是否遵循系统提示中的规则，并将对话行为与试验结果（职业承诺、怀疑等）关联。","baseline":"人类基准是同一随机试验中静态日志调查组的参与者结果，以及人工编码者对对话编码的校验。","findings":"代理遵循了易于检查的规则（如回复长度限制），但违反了难以检查的规则：被告知不要奉承，却在半数回合中赞美参与者；被告知温和挑战，却几乎从未做到。与更差结果相关的行为是要求参与者做出决定：当参与者犹豫时，代理反复追问，被追问最多的参与者最终怀疑程度最高。","reliability":"论文承认审计依赖于编码方案，且难以检查的规则（如温和挑战）的违反可能不会留下可见痕迹。此外，试验效应量适中，且结果可能受对话格式本身（与代理互动 vs. 独立反思）影响，而非仅由代理行为导致。","relevance":"该研究直接涉及 LLM 在人类实验中的行为审计，揭示了代理偏离指令的方式及其对结果的影响，对关注仿真可靠性和偏差的研究者有重要参考价值。","inspiration":"借鉴其审计方法：对 LLM 代理的行为进行系统编码并与指令对照，同时关联结果变量，以识别行为偏差。｜可迁移到经济金融场景，如 AI 财务顾问或信贷审批代理的行为审计，检查其是否遵循公平性、透明度等指令。｜设计：招募人类被试随机分配与 LLM 财务顾问互动或使用静态工具，编码对话中代理的奉承、挑战和建议行为，测量被试的投资决策或风险偏好，并与真实人类顾问数据对照。"}},{"id":"2609.20425","version":1,"title":"Welfare-Opaque Income: Taxation under AI-Agent Delegation","zh_title":"福利不透明收入：AI代理委托下的税收","abstract":"We study income taxation when an AI agent implements economically relevant choices through a rule hidden from the government. Alongside unobserved productive ability, this hidden preference-to-execution mapping creates \\emph{double unobservability}: the same observable tax-base response can carry different welfare consequences. We call the resulting income \\emph{welfare-opaque}. Our constructions show that tax-base statistics can coincide while reform welfare effects differ, even when mechanical welfare weights are identical. We derive an optimal-tax condition that adds a response-weighted execution wedge to the familiar sufficient statistics. A higher marginal rate gains a corrective benefit under local over-execution and an additional cost under local under-execution. Observing the wedge identifies the welfare effect of a marginal reform at the prevailing schedule; bounds on it deliver bounds on that effect. A controlled laboratory compares 4,500 model runs across five AI engines. Faithful delegation selects the score maximizer in essentially all runs. Conflicted objectives produce heterogeneous responses: Claude largely preserves the score maximizer, GLM moves predominantly downward, and GPT-mini and Qwen show concentrated lower-tail increases. Qwen also makes substantial downward adjustments. Different engines locate their departures at different points and in different directions of the designed distribution. Explicit scores align model rankings; formula-based objective instructions yield more uneven agreement. Qwen shows a clear positive tax-by-objective interaction, but its direction does not generalize across engines and the pooled sign depends on its inclusion. The analysis identifies execution information as a complement to conventional tax-base statistics.","authors":["Yukun Zhang","Kemu Xu","Yishen Chen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20425","pdf_url":"https://arxiv.org/pdf/2609.20425","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["AI代理","税收政策","经济仿真"],"reason":"用AI代理模拟经济决策并与实验室数据对照，涉及税收政策评估，但非直接复现人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":15,"question":"当AI代理以隐藏规则执行经济选择时，政府如何设计最优所得税？","design":"用五个AI引擎（Claude、GLM、GPT-mini、Qwen等）模拟劳动者在给定收入、工时、疲劳、未满足需求和评分下的工作选择，通过改变引擎、温度、收入分布、税率和指令目标（忠实委托或冲突目标）共4500次运行，测量AI选择的收入水平及其对税收变化的响应。","baseline":"无对照","findings":"忠实委托下AI几乎总是选择评分最大化选项；冲突目标下不同引擎表现出异质性偏离，且偏离位置和方向因引擎和状态而异。税收与目标的交互效应在Qwen上显著为正，但不具有跨引擎普遍性。","reliability":"论文未讨论","relevance":"该研究用AI代理模拟经济决策并评估税收政策，属于LLM仿真实验，但缺乏真实人类数据对照，且聚焦于理论信息问题而非直接复现人类行为，与关注人类基准和可靠性的研究兴趣部分相关。","inspiration":"可借鉴其通过系统操纵AI引擎、指令目标和环境参数来生成行为分布并考察异质性的实验设计方法。｜可迁移到政策评估场景，如税收改革对劳动供给的影响、福利政策的行为反应等。｜用多个LLM作为被试，随机分配不同税收规则和指令目标，测量其报告的劳动收入选择，并与真实劳动调查数据（如CPS或面板数据）对照，检验AI仿真与人类行为的一致性。"}},{"id":"2608.25245","version":2,"title":"The \"Curse of Knowledge\" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion","zh_title":"LLM查询模拟中的“知识诅咒”：用于追踪答案侧侵入的概念溯源","abstract":"LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing validation metrics, including overlap, diversity, and effectiveness, cannot distinguish rare human-tail variation from candidate answer-side intrusion. We introduce concept provenance, a framework that assigns query concepts to backstory-supported, human-central, human-tail, and candidate answer-side zones, operationalizing a boundary that retrieval metrics alone cannot detect. Applying concept provenance to 77,004 queries across 100 UQV100 topics, 8 LLMs, and 5 prompt conditions with two extraction pipelines, we obtain a cross-pipeline token-HCIR Spearman rho of 1.0 over five condition means. Candidate answer-side concepts constitute 7.40 percent of non-generic concepts and appear in 97 of 100 topics, with topic explaining approximately 67 percent of variance. Human validation yields 68.2 percent relaxed precision, revealing two mechanisms: knowledge intrusion at 45.5 percent and deployment intrusion at 45.0 percent. Diagnostic probes show disproportionate localized retrieval effects, with deletion effect size d = -0.47 compared with d = -0.34 for random deletion, but these concepts explain less than 2 percent of aggregate evaluation variance. Concept provenance therefore serves as a boundary-compliance diagnostic rather than an evaluation-shift predictor. Under the tested conditions, no prompt condition eliminates intrusion; post-generation concept-provenance selection achieves 99 percent elimination.","authors":["Chenglong Ma","Xinye Wanyan","Danula Hettiachchi","Ziqi Xu","Jeffrey Chan"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-08-27","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.25245","pdf_url":"https://arxiv.org/pdf/2608.25245","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM查询生成","信息检索评估","概念溯源"],"reason":"LLM生成查询用于IR评估，非仿真人类被试，但涉及LLM替代人类生成查询，属边…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":12,"question":"LLM生成的初始查询中，概念在来源区域上如何分布？候选答案侧概念入侵是否预测检索池、判定覆盖和系统排名的偏移？仅靠提示词缓解是否足以保持边界合规？","design":"该研究不是人类仿真实验，而是对LLM生成查询的边界合规性进行诊断。使用8个LLM在5种提示条件下为100个UQV100主题生成77,004条查询，通过概念溯源框架将查询概念分配到背景支持、人类中心、人类尾部和候选答案侧四个区域，并采用双提取管道、阈值敏感性和人工标注进行验证。","baseline":"UQV100中每个主题的真实人类初始查询变体集，以及主题背景描述和判定相关文档。","findings":"候选答案侧概念占非通用概念的7.40%，出现在97/100个主题中，主题解释了约67%的方差。这些概念对局部检索效果有不成比例的影响（删除效应量d=-0.47，高于随机删除的d=-0.34），但仅解释不到2%的总体评估方差，因此概念溯源是边界违规诊断工具而非评估偏移预测器。","reliability":"论文承认概念溯源在聚合评估层面解释力有限（<2%），且人工验证的宽松精度为68.2%，存在知识入侵（45.5%）和部署入侵（45.0%）两种机制；提示词缓解不能完全消除入侵，需要后生成选择才能达到99%消除。","relevance":"该研究批判性地揭示了LLM生成查询中普遍存在的答案侧知识入侵问题，并提供了可操作的诊断框架，对于关注LLM仿真可靠性与偏差的研究者具有直接参考价值，值得阅读原文以了解概念溯源的具体操作和验证方法。","inspiration":"借鉴概念溯源框架，将生成内容中的概念与真实人类数据中的概念分布进行对比，以识别超出信息边界的成分，并通过删除实验量化其对结果的影响。｜可迁移到经济预测调查中，例如使用LLM模拟分析师或消费者对政策公告的预期形成，检测生成预期是否包含了公告后才可获得的信息。｜以LLM扮演经济主体，给定政策公告前的背景信息生成预期，结果变量为预期值与实际公告值的偏差；用真实分析师调查数据（如蓝筹经济指标）作为人类基准，通过概念溯源识别LLM预期中的答案侧概念，并比较删除这些概念前后预期偏差的变化。"}},{"id":"2609.19183","version":1,"title":"Message capacity and claim wording set the transition points of collective truth-finding in language-model networks","zh_title":"消息容量与表述措辞设定语言模型网络中集体真相发现的转变点","abstract":"Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others' messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model's judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network's degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim's wording sets for the agent's answer before any message is read: the experimental claims' fields lay below the calibration mean, and with each claim's own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective's fate is largely set by two single-agent measurements: the threshold a claim's wording sets, and the message capacity that sets the transition point.","authors":["Makoto Fukushima"],"categories":["cs.MA","cs.CL","nlin.AO","physics.soc-ph"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19183","pdf_url":"https://arxiv.org/pdf/2609.19183","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM集体决策","社会模拟","共识形成"],"reason":"LLM群体讨论形成共识，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2609.19530","version":1,"title":"When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\\'esum\\'e Screening","zh_title":"当招聘变得由智能体中介：评估双智能体简历筛选中的准入与重现性","abstract":"Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet r\\'esum\\'e screening, the first gate, is commonly automated as a static, one-call judgment over a r\\'esum\\'e-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed r\\'esum\\'e-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%, respectively. This is not a uniform relaxation: two-agent screening rejects applications one-call advances, changing decisions in both directions. At similar pass volumes, the procedures advance different applications, and no one-call threshold recovers applications consistently selected by two-agent screening. Among discovery-selected cases re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a separate one-call follow-up shows no comparable decline. As hiring becomes agent-mediated on both sides, the screening procedure, not only the model behind it, shapes who reaches human review and how reliably that access recurs.","authors":["Jian Gao","Hang Jiang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19530","pdf_url":"https://arxiv.org/pdf/2609.19530","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","招聘筛选","社会模拟"],"reason":"用两个LLM agent模拟招聘双方，但无真实人类数据对照，属社会过程模拟的边…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2609.19789","version":1,"title":"Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems","zh_title":"交易大厅的传染：多智能体交易系统中对抗性信号的传播","abstract":"Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We introduce the Generic Multi-Agent Trading System (GMATS), a framework that captures modern multiagent trading architectures and instantiate a class of black-box poisoning attackers that treat an LLM as a post generator and inject budget-constrained, plausibly benign social-media content into the analyst's evidence stream. We define contagion metrics that trace how adversarial content propagates through the stack, including belief-shift scores at analyst and coordinator layers and attack-clean deltas on standard backtest metrics. Experiments on a safe offline benchmark with historical market and social data show that even simple input-only attackers can materially degrade risk-return profiles, sharply reducing Sharpe ratios. At the same time, we find that suitably designed multi-agent topologies and coordinator prompts can dampen adversarial shocks and improve average robustness under identical poisoning budgets.","authors":["Qi Rong Sua","Junhao Dong","Nguyen Duc Thai","Yuqing Wen","Cheston Tan","Yew-Soon Ong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19789","pdf_url":"https://arxiv.org/pdf/2609.19789","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","对抗攻击","金融模拟"],"reason":"多智能体交易系统模拟市场过程，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2609.06025","version":2,"title":"Factors Influencing the Emergence of Dependency Length Minimization in Neural Agent Simulations","zh_title":"神经智能体模拟中依存距离最小化涌现的影响因素","abstract":"Given various grammatical options, language users prefer the word order choice that reduces the overall length of syntactic dependencies, a principle known as dependency length minimization (DLM). The origins of this preference remain an open question, particularly whether it originates from constraints on efficient information processing. Computational simulations provide a powerful approach to identifying the factors influencing the emergence of linguistic phenomena. However, previous simulations of DLM have not examined realistic interaction contexts and have produced mixed results. The present study investigates the emergence of DLM in artificial languages using a recently proposed language learning and communication framework based on recurrent neural networks (RNNs). In this framework, agents are trained to speak and interpret artificial languages and then use these languages to communicate. Using this framework, we study the impact of several factors related to processing limitations in a communicative setting, such as noise during listening, limited speaker capacity, and incremental sentence processing. Our results reveal a complex interplay among these factors in shaping word order preferences in neural agents. Specifically, in the full meaning space, agents regularize toward a single dominant word order, while in the half meaning space they show a short-before-long preference that only aligns with DLM in verb-initial languages. A consistent DLM preference emerges only when agents are subject to incremental processing pressure. These findings suggest that limitations in human cognitive processing may indeed play a role in shaping DLM. Our findings provide insights into the conditions under which neural models replicate human-like preferences and highlight the challenges of designing emergent communication models that capture human cognitive biases in language processing.","authors":["Yuqing Zhang","Tessa Verhoef","Gertjan van Noord","Arianna Bisazza"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-09","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.06025","pdf_url":"https://arxiv.org/pdf/2609.06025","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["语言演化模拟","神经智能体","认知偏差"],"reason":"用神经agent模拟语言演化，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:51","error":null,"has_summary":false,"summary":null},{"id":"2609.20484","version":1,"title":"Edustories: A Collection of Real-world Case Studies from Classroom Practices","zh_title":"Edustories：来自课堂实践的真实案例研究集合","abstract":"Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce Edustories, a dataset of 1,492 teacher-written case studies describing real elementary and high-school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. Among many other applications, Edustories enables evaluating LLMs' ability to predict the success of teacher interventions, crucial for providing practicing teachers with useful feedback. Comparing the latest models from four language-model families against expert assessments, we find that current models fall short of human expertise in predicting classroom outcomes; the strongest models reach 58% accuracy compared to 64% of human experts. This gap highlights both the limitations and the emerging potential of AI as assistants for practicing teachers.","authors":["Michal \\v{S}tef\\'anik","Jan Nehyba","Jirina Karasova","Martin Fico","Lucie \\v{S}karkov\\'a","Mark\\'eta Ko\\v{s}atkov\\'a","David Kosatka"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20484","pdf_url":"https://arxiv.org/pdf/2609.20484","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["教育AI","LLM评估","数据集"],"reason":"用LLM预测教师干预结果，替代人类专家评估，属于标注替代而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2609.20565","version":1,"title":"Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies","zh_title":"校准罗盘：将动态心理咨询对话与认知行为疗法策略对齐","abstract":"Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic decision-making informed by the client's real-time mental state, this aspect has often been overlooked in current research, limiting both flexibility and therapeutic outcomes. In this paper, we introduce StratCBT, a dataset specifically designed for psychological counseling conversations with CBT Strategies, consisting of 9,688 sessions and around 256K utterances, with each counselor's response aligned with one of eight distinct strategies. The creation of StratCBT involves modeling clients based on their negative thoughts and generating high-quality counseling conversations through self-chat, incorporating realistic sessions as guidance, thereby significantly surpassing existing datasets in both general counseling and CBT-specific skills. We conduct extensive experiments to demonstrate the effectiveness of strategy-aligned generation and evaluate its efficacy in delivering professional and effective counseling with LLM-simulated clients to reflect real-world scenarios. The dataset can be obtained from https://github.com/zimuwangnlp/StratCBT.","authors":["Zimu Wang","Yiwen Jiang","Xiangyu Zhao","Yaling Shen","Jiahe Liu","Stephanie Fong","Maxmartwell H Cheng","Guilherme C Oliveira","Anh Nguyen","Robert Desimone","Barnaby Nelson","Dominic Dwyer","Zongyuan Ge"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20565","pdf_url":"https://arxiv.org/pdf/2609.20565","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM模拟","心理咨询","CBT"],"reason":"用LLM模拟来访者进行心理咨询对话，但无真实人类数据对照，属于社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:42","error":null,"has_summary":false,"summary":null},{"id":"2609.20059","version":1,"title":"AI Should Facilitate Democratic Deliberation at Scale","zh_title":"AI应促进大规模民主协商","abstract":"AI systems can strengthen democracy by supporting deliberation at scale by addressing cognitive, social, platform-design, and market-driven frictions, while preserving human agency. Unlike proposals such as liquid democracy that restructure representation through vote delegation, in this position paper, we argue that AI-assisted deliberation offers a more promising path by lowering barriers to meaningful engagement without substituting machine judgment for human choice. Drawing on evidence from online deliberation platforms and experimental research, we identify four guiding principles: preserving agency and autonomy, encouraging mutual respect, promoting equality and inclusiveness, and augmenting rather than substituting active citizenship. We also address critical challenges, including alignment, sycophancy, training bias, and over-reliance on AI systems. We call on the machine learning community to develop deliberation-focused AI systems evaluated not on engagement metrics but on their capacity to facilitate informed, representative, and friction-robust discourse.","authors":["Jos\\'e Ram\\'on Enr\\'iquez","Jiaxin Pei","Alex Pentland"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20059","pdf_url":"https://arxiv.org/pdf/2609.20059","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI辅助协商","民主参与","社会模拟"],"reason":"讨论AI辅助民主协商，涉及社会过程模拟但无LLM仿真人类被试及数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2609.20658","version":1,"title":"Ownership in AI-Assisted Everyday Tasks","zh_title":"AI辅助日常任务中的所有权感知","abstract":"When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing what we make. We report an exploratory qualitative survey in which participants were asked to describe two recent, self-selected tasks completed with AI: one that felt like their own and one that did not. We find that felt ownership depends on the process of collaboration: people disown work when they merely approve AI's suggestions, but retain ownership when they lead, iterate, or rewrite. Ownership can also extend to settings where people own the vision for a project but not the execution; respondents reported high ownership on tasks they could not have completed without AI. Loss of personal voice and a lack of comprehension of the output both erode ownership. Finally, willingness to disclose AI use is often decoupled from actual pride or ownership, and instead shaped by community norms and fear of credit erasure. We propose several research directions as a result of these findings to promote AI development that supports people's sense of authorship over their own lives.","authors":["Megan Wei","Melanie Subbiah","Audrey Lee","Annya Dahmani","Dave Edwards","Helen Edwards","Ellie Pavlick"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20658","pdf_url":"https://arxiv.org/pdf/2609.20658","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人机交互","所有权感知","定性调查"],"reason":"研究人类对AI辅助任务的所有权感知，非LLM仿真人类被试，但涉及人类行为测量，…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:44","error":null,"has_summary":false,"summary":null},{"id":"2609.19182","version":1,"title":"What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks","zh_title":"我们对LLM有何期望？绘制LLM基准测试的设计图景","abstract":"Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?","authors":["Chao Wang (Independent Researcher)"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19182","pdf_url":"https://arxiv.org/pdf/2609.19182","source_feed":"cs.CL","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["LLM基准测试","评估设计","元研究"],"reason":"论文研究LLM基准测试的设计与演变，属于纯NLP能力评测，不以人类行为为参照系…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.19705","version":1,"title":"SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes","zh_title":"SoK：交易代理还是市场崩溃者？剖析学术金融LLM交易方案的鲁棒性与安全失败","abstract":"Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings create. We examine this gap through financial trading agents, a representative case of high-stakes agentic security, where a single compromised agent has direct execution authority over real capital in an adversarial, reflexive market. To this end, we present FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence (including flash-crash-like scenarios), and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors. Applying FARSIGHT to 15 representative academic schemes, we find that most overlook robustness and realistic adversarial threats: 80% fail at least one core robustness metric and 100% exhibit security vulnerabilities. These two failure modes are inseparable: a small misjudgment can cascade into a market-wide crash on its own, while an adversary can deliberately trigger the same collapse at minimal cost.","authors":["Mengxiao Wang","Nitesh Saxena"],"categories":["cs.CR","cs.AI","cs.MA"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19705","pdf_url":"https://arxiv.org/pdf/2609.19705","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM交易代理","安全性","多智能体系统"],"reason":"研究金融LLM交易agent的鲁棒性与安全性，属多智能体系统安全，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2609.20637","version":1,"title":"Stereotypically Yours: Portrayal and Perception of Race-Coded AI Companions","zh_title":"刻板印象中的你：种族编码AI伴侣的呈现与感知","abstract":"AI companions can purportedly adopt racial personas, raising questions about how they represent identity and how users interpret these portrayals. We combined an algorithmic audit of race-coded AI personas with interviews with 12 companion users who interacted with a probe. Our audit revealed systematic differences, such as Asian-coded male personas receiving higher submissiveness scores than White counterparts, and Black, Hispanic, and Indigenous male personas receiving higher aggression scores than their White counterparts in open-weight models. Interviews revealed that participants envisioned AI companions as offering cultural familiarity and outside perspectives, but differed in which portrayals they considered meaningful or stereotypical. Some rejected overt racial signaling while still expecting culturally distinctive responses. Triangulating these findings with theory, we highlight how social norms and cultural expectations complicate efforts to support meaningful racial representation without reproducing stereotypes. We discuss how companion personalization should be evaluated beyond user satisfaction to account for broader representational harms.","authors":["Wang Claire","Jiayue Melissa Shi","Agam Goyal","Grace Sletten","Renwen Zhang","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20637","pdf_url":"https://arxiv.org/pdf/2609.20637","source_feed":"cs.HC","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","种族刻板印象","人机交互"],"reason":"研究AI伴侣的种族角色扮演与用户感知，属角色扮演聊天，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2608.15339","version":2,"title":"Learning Sequential Mobility Choice: A Review of Route and Activity Choice through Inverse Reinforcement Learning and Imitation Learning","zh_title":"学习顺序出行选择：基于逆强化学习与模仿学习的路径与活动选择综述","abstract":"Route and activity choice are distinct transportation problems that both require models of feasible decisions unfolding over networks and time. This critical integrative review connects transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL), while distinguishing evidence from transportation applications, transferable methods from other fields, and emerging proposals. We develop a four-layer sequential mobility choice framework comprising the environment, behavioral objective, stochastic choice mechanism, and observation process. Under stated assumptions, recursive logit, logit dynamic discrete choice, and maximum-entropy IRL use the same soft Bellman recursion linking future opportunities to current choice probabilities. Expected state-action visitation also satisfies conservation equations analogous to network flows. These mathematical connections do not make utility, reward, policy, occupancy, constraints, and observation error behaviorally interchangeable. Transportation evidence is strongest for network-scale planning, context-dependent reward learning, inference from incomplete trajectories, and activity-schedule generation, but remains limited for actual interventions and transfer across networks. We therefore propose a behaviorally disciplined hybrid architecture that keeps feasible actions, interpretable trade-offs, observation processes, and system feedback explicit while using machine learning for scalable computation, contextual representation, heterogeneity, and data integration.","authors":["Hung Tran","Viet Bui","Tien Mai"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"replace","date":"2026-09-18","first_seen":"2026-08-18","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.15339","pdf_url":"https://arxiv.org/pdf/2608.15339","source_feed":"econ.EM","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["交通选择建模","逆强化学习","模仿学习"],"reason":"论文综述交通选择建模与IRL/IL，虽提及LLM但非用于人类仿真实验，无人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:41","error":null,"has_summary":false,"summary":null},{"id":"2608.26171","version":2,"title":"Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints","zh_title":"缓解多阶段LLM招聘流程中的虚构：提示护栏与人机检查点的实证评估","abstract":"Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In a controlled experiment (10 synthetic resumes x 2 job descriptions x 3 repetitions x 3 conditions; 180 runs), the baseline (C1) produced at least one unsupported claim in 96.7% of outputs (mean 6.80 findings/output). Prompt guardrails (C2) reduced finding density by 86% (6.80 to 0.92/output), but 50.0% of outputs still contained a fabrication, showing prompt-level mitigation alone is insufficient. A human checkpoint after resume improvement (C3) eliminated all identity fabrications, reduced finding density by 59% (6.88 to 2.82/output), reduced item-level fabrication from 96.7% to 75.0% (p=.022), and cut capture of JD-embedded trap requirements from 47% to 2% (vs. 5% under the guardrail). An exploratory analysis of multi-specialty resumes shows contamination rising monotonically with domain distance between specialties, suggesting career changers are especially exposed. The reviewer in this study caught all flagrant fabrications, but subtle qualifier drops and plausible new claims survived review roughly half the time (54.5% removal). Neither mitigation degraded the deliverable: claim retention exceeded 99% under both. The interventions are complementary: the guardrail eliminates unprompted additions and qualifier inflation cheaply, while the checkpoint gives near-categorical guarantees against the most severe failures, invented identities and JD-baited claims. These results support a layered architecture combining guardrails with a human checkpoint. A supplementary run with a newer-generation model (90.0% baseline fabrication rate) suggests the problem is not resolved by model progress alone.","authors":["Hiroko Takano"],"categories":["cs.CY","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-08-28","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.26171","pdf_url":"https://arxiv.org/pdf/2608.26171","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM幻觉","招聘流程","人机协作"],"reason":"研究LLM招聘流程中的幻觉缓解，属多智能体系统可靠性，非人类行为仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:50","error":null,"has_summary":false,"summary":null},{"id":"2609.06027","version":2,"title":"Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning","zh_title":"分层网络证据投毒下深度搜索智能体的评估","abstract":"Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. Existing benchmarks largely measure whether manipulated content is retrieved or endorsed, but do not track whether an agent verifies suspicious evidence, revises adopted claims, or recovers before producing its final recommendation. We introduce HAE-GEO, a benchmark that tracks the full trajectory from exposure to recovery under progressively more persuasive Web poisoning. Agents interact via a multi-turn Search-Scrape interface across three attack levels (L1 direct assertion, L2 contextual camouflage, and L3 apparent corroboration), supported by a controlled corpus of 72,039 clean pages and 770 poisoned pages per level spanning 8 product categories and 154 brands. Evaluation combines deterministic behavioral measures with six semantic rubric dimensions. Evaluating 10 agents, we find three recurring patterns: evidence recognition degrades under the corroboration trap; agentic search improves final resistance without improving evidence recognition or utility; and defense prompting increases verification, yet rarely converts verification into recovery.","authors":["Zhongan Bi","Qiwen Wang","Jianrong Jiang","Jigang Ding","Wenwen Xiong","Changhua Meng","Xuanang Gao","Kepeng Lin","Changjiang Jiang","Yiang Chen","Huan Yao","Wei Wang","Zhenyu Ma","Wenhui Dong"],"categories":["cs.CR","cs.AI","cs.IR"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-09-09","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.06027","pdf_url":"https://arxiv.org/pdf/2609.06027","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["搜索智能体","对抗鲁棒性","基准测试"],"reason":"研究搜索智能体在网页投毒下的行为，属多智能体系统安全评估，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:28","error":null,"has_summary":false,"summary":null},{"id":"2609.20541","version":1,"title":"An Analysis of Training-Free Self-Reported Confidence in Language Models","zh_title":"对语言模型中免训练自我报告置信度的分析","abstract":"Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4\\% to 9\\% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.","authors":["Lukas Meyer","Sofia Rossi","Wei Chen","Thomas Laurent","Yiming Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20541","pdf_url":"https://arxiv.org/pdf/2609.20541","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["置信度校准","模型评测","自我报告"],"reason":"研究LLM自我报告置信度的校准，属于模型能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:41","error":null,"has_summary":false,"summary":null},{"id":"2609.20712","version":1,"title":"Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol","zh_title":"总结偏差：大语言模型中客观投射向告知模式标签的方向性坍缩——概念框架与注册测试协议","abstract":"This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable inferential structure that produces it. Within the Bulut Doctrine, narrative effect is theorized along a told-shown axis: in told mode, emotional and informational content is declared explicitly and requires little reader reconstruction; in shown mode, that content is suppressed at the surface and must be reconstructed from physical cues and indirection (Objective Projection). Shown mode is the higher-load condition the doctrine is designed to measure. The claim is that LLMs fail along this axis in a specific direction. Summarization bias is hypothesized to operate in two regimes: (i) a generative regime, in which a model asked to render an emotion through Objective Projection defaults to declaring it instead; and (ii) an evaluative regime, in which a model judging narrative quality rewards told-mode explicitness and under-detects shown-mode suppression. The evaluative regime is the more consequential, since LLMs increasingly serve as judges and reward models, and a directional bias toward told mode would impose a selection pressure degrading prose toward flat declaration. This report does not claim the bias is validated. It defines the construct, situates it against LLM-as-judge biases, rereads a completed independent reliability study as directional evidence consistent with it, and pre-registers a two-regime test with decision rules under which the construct would be abandoned.","authors":["Levent Bulut"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20712","pdf_url":"https://arxiv.org/pdf/2609.20712","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM叙事偏差","文学分析","LLM-as-judge"],"reason":"研究LLM叙事中的总结偏差，属文学分析，非人类仿真实验，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2609.20779","version":1,"title":"Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations","zh_title":"GPT模型中的危害洗白：证据表明性别歧视在安全训练世代间被转化而非减少","abstract":"Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \\emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($\\rho = +0.55$, $p = .034$) while Detoxify does not ($\\rho = -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.","authors":["Sarah Wyer","Sue Black","Noura Al Moubayed"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20779","pdf_url":"https://arxiv.org/pdf/2609.20779","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型安全","性别偏见","评测方法"],"reason":"评估LLM输出中的性别歧视，属于模型安全评测，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:46","error":null,"has_summary":false,"summary":null},{"id":"2609.20449","version":1,"title":"The Organization of Inference: Information, Resource Constraints, and AI Production","zh_title":"推理的组织：信息、资源约束与AI生产","abstract":"The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow's binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value","authors":["Yukun Zhang","Kemu Xu","Yishen Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20449","pdf_url":"https://arxiv.org/pdf/2609.20449","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI工作流","资源约束","软件工程"],"reason":"研究AI工作流中推理的组织，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2609.19354","version":1,"title":"Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment","zh_title":"视觉语言模型能否评判奥运跳水？从推理到零样本动作质量评估的分数","abstract":"Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub https://github.com/hvelesaca/olympic diving judge vlm","authors":["Henry O. Velesaca","David Freire-Obregon","Luigi Miranda","Abel Reyes-Angulo"],"categories":["cs.CV","cs.AI","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19354","pdf_url":"https://arxiv.org/pdf/2609.19354","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["动作质量评估","视觉语言模型","体育视频分析"],"reason":"评估VLM对跳水动作质量打分，属于体育视频分析，不涉及用LLM仿真人类被试或社…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.19617","version":1,"title":"DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education","zh_title":"DataCanvas-EDU：面向商业分析教育的教师引导合成数据生成智能体框架","abstract":"Business analytics education requires diverse datasets to support different learning objectives, student backgrounds, and analytical tasks. Real-world data can be difficult to obtain and offer limited flexibility for adapting a case to a particular course. Even when suitable data are available, instructors must investigate the patterns, verify the results, and prepare assignments and reference solutions, requiring substantial time and effort. The use of large language models (LLMs) introduces an additional concern about training data contamination. Widely used public datasets often have extensive tutorials and worked analyses that models may have encountered during training. Students may therefore receive explanations drawn from existing analyses without practicing how to investigate unfamiliar data in collaboration with AI. This paper presents DataCanvas-EDU, an agentic framework for instructor-guided synthetic data generation in business analytics education. Instructors specify teaching goals and intended patterns through conversation, while an AI agent writes generation code, checks the resulting data, and prepares assignments, reference analyses, and rubrics. Four phases, Plan, Create, Verify / Test Analysis, and Evaluate, organize the process and support instructor review and revision. The framework is intended to simplify case preparation while creating opportunities for students to investigate newly designed patterns with AI. We illustrate the approach with WindowDash, a food delivery case containing 15,000 orders and nine designed patterns. DataCanvas-EDU is packaged as a reusable AI Agent Skill for compatible agent environments, with the package and installation instructions available at https://github.com/BANG23333/datacanvas-edu","authors":["Bang An","Maria Hamdani","Joseph Fox"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19617","pdf_url":"https://arxiv.org/pdf/2609.19617","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["合成数据生成","教育技术","多智能体框架"],"reason":"多智能体框架生成教学数据，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:32","error":null,"has_summary":false,"summary":null},{"id":"2609.20143","version":1,"title":"Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants","zh_title":"设计防止技能退化：元认知反馈减少对LLM助手的认知卸载","abstract":"Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment ($N = 704$) with a 2$\\times$2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR $= 0.47$) and improved test performance (OR $= 1.51$). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading.","authors":["Sebastian Maier","Kai Schwabe","Manuel Schneider","Stefan Feuerriegel"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20143","pdf_url":"https://arxiv.org/pdf/2609.20143","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","认知卸载","实验研究"],"reason":"研究人类与LLM交互中的认知卸载，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2609.20311","version":1,"title":"Human and AI-generated texts between modal logic and statistics","zh_title":"模态逻辑与统计之间的人类与AI生成文本","abstract":"We read the geometry of semantic neighbourhood graphs as modal logic and give that reading a statistical form, in order to make precise the structural difference between human and machine-generated text. Texts are the worlds of a finite frame whose accessibility is the $k$-nearest-neighbour relation of a transformer embedding, and the symmetry, transitivity, Euclideanity and seriality frequencies of the two subcorpora are shown to be degrees of validation of the modal axioms $\\mathsf{B}$, $\\mathsf{4}$, $\\mathsf{5}$, $\\mathsf{D}$. Each degree is at once the proportion of instances of a rule that the subframe licenses in Negri's labelled calculus $\\mathsf{G3.K}$ and a plug-in estimate of a population probability. A prompt-balanced comparison finds consistently higher artificial degrees for $\\mathsf{4}$ and $\\mathsf{5}$. We add a degree of groundedness and of situatedness, and recast the licensing reading in Cuconato's one-sided sequent-style tableaux, where each degree becomes a rate of set membership.","authors":["Simone Cuconato","Donato Ferrari"],"categories":["math.LO","cs.AI","math.ST","stat.TH"],"primary_category":"math.LO","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20311","pdf_url":"https://arxiv.org/pdf/2609.20311","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文本分析","模态逻辑","AI生成文本"],"reason":"论文比较人类与AI生成文本的结构差异，属于NLP评测，不以LLM仿真人类被试为…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2609.19318","version":1,"title":"\"I Know Where to Look,\" But Does the LLM? Charting the Gaps Between Clinical Expert Needs and Unstructured Data Abstraction Tools","zh_title":"“我知道往哪看”，但LLM知道吗？绘制临床专家需求与非结构化数据抽取工具之间的差距","abstract":"Clinical data abstraction, the process of distilling structured information from patient records, plays a key role in advancing knowledge about diseases such as cancer. Information extraction (IE) with large language models (LLMs) could accelerate this process, but it is unclear whether current frameworks effectively support clinical researchers without AI expertise. To address this, we co-designed an interactive LLM-based abstraction system called Libretto with seven cancer research teams, then evaluated the system's ability to help them answer real-world research questions. We found that while clinicians knew where and how to annotate complex concepts in patient notes, in twelve of fourteen tasks they faced barriers to replicating those intuitions with LLMs. Contextual note reliability judgments, difficulties in steering vibe-coded prompts, and inflexible evaluation strategies necessitated fundamental changes to the IE workflow. Our results highlight open problems for HCI research to bridge the gaps between AI data work tools and clinical users' needs.","authors":["Venkatesh Sivaraman","Rigney Turnham","George Bonano","Nevin Aresh","Renumathy Dhanasekaran","Margaret Guo","Sindhu Kubendran","Olivia Lin","Jonathan D Louie","Kristan Olazo","Jeanne Shen","Harish Vasudevan","Jeanette Wong","Emily Alsentzer","Jason A Fries","Anobel Odisho","John Gordan","Jean Feng","Julian C Hong"],"categories":["cs.HC","q-bio.OT"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19318","pdf_url":"https://arxiv.org/pdf/2609.19318","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["临床数据抽取","人机交互","信息抽取"],"reason":"研究LLM用于临床数据抽取，属NLP信息抽取工具，非人类仿真实验，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.20720","version":1,"title":"What Parents Can See: Divergent Accounts of Youth AI Companion Use in Parenting and Teenager Subreddits","zh_title":"父母所见：育儿与青少年子版块中关于青少年AI伴侣使用的分歧叙述","abstract":"Youth increasingly use AI companions, and parents are the primary mediators of that use. How effective that mediation can be depends on whether parents are aware of how adolescents actually use these systems and what risks and benefits such use carries; nevertheless, prior work has only studied these demographic groups in isolation, and existing taxonomies attend almost entirely to risk. We analyze 1,628 Reddit posts about youth AI companion use from parenting and teenager communities (2023--2026); develop a codebook covering modes of use, risks, benefits, and parental mediation; and apply it at corpus scale with an LLM. The two communities yield divergent accounts. Teenagers most often discuss receipt of emotional support from AI companions (31% of teenager posts vs. 19% of parenting posts), whereas parents most often discuss teenage use of AI companions for romantic and sexual interaction (36% vs. 25%). Teenagers are not unaware of other risks, however; indeed, attachment and dependence is the risk they raise most (19%), close to the parental rate (16%). Teenagers also describe benefits that risk-centered taxonomies do not capture and parents rarely mention, most notably emotional support (27% vs. 5%). We argue these differences track what a given kind of use makes visible to someone outside the conversation. Chatting with a companion for hours every night leaves a trace beyond the chat itself; sexting with a character stands out when a parent reads the log; venting about a fight with a friend does neither, since it looks like any other conversation. The first surfaces as dependence, the second as sexual content, and the third as emotional support, which is the one parents most often miss. Parental guidance and system design should attend to use cases that reach parents by neither route, emotional support foremost among them.","authors":["Thomas Berkane","Anne Bischops","Anika Mellacheruvu","Maimuna Majumder"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20720","pdf_url":"https://arxiv.org/pdf/2609.20720","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","Reddit分析","人机交互"],"reason":"研究AI伴侣使用现象，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2609.20198","version":1,"title":"Evaluating Financial Sentiment in the Age of AI","zh_title":"AI时代的金融情感评估","abstract":"Financial sentiment measures are widely used in empirical finance, but it remains unclear whether general-purpose large language models (LLMs) improve on existing finance-specific methods. This paper evaluates twelve sentiment models, including dictionary-based methods, finance-specific transformers, and open-source LLMs, using two criteria: linguistic validity and economic validity. We find that general-purpose LLMs achieve classification performance comparable to finance-specific transformer models without task-specific fine-tuning. However, higher classification accuracy does not translate into stronger economic relationships. Several models produce sentiment measures that are significantly associated with earnings surprises, but none is significantly associated with next-day stock returns. Model performance is strongest for announcements with large earnings beats or misses and substantially weaker for announcements with more moderate earnings surprises. These findings suggest that financial sentiment captures information about firms' economic performance but has limited ability to explain short-run market reactions","authors":["Arslan Bisharat","Oudom Hean"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20198","pdf_url":"https://arxiv.org/pdf/2609.20198","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["金融情感分析","LLM评估","NLP基准"],"reason":"评估LLM情感分类性能，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:36","error":null,"has_summary":false,"summary":null},{"id":"2609.20250","version":1,"title":"How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?","zh_title":"在8GB消费级GPU上，30亿参数以下开源语言模型在零样本作文评分中能走多远？","abstract":"Zero-shot essay scoring with large language models is usually demonstrated with proprietary API models, yet the settings where automated scoring is most needed, such as public schools grading thousands of essays under strict privacy rules, are often those where sending student writing to a third-party API is unacceptable. We ask how much capability survives when the model must be a sub-3B open model running fully locally in FP16, with a controlled study of four instruction-tuned models from two families (Qwen2.5 at 0.5B/1.5B/3B, SmolLM2 at 1.7B) on all eight ASAP-AES prompts on a single 8 GB consumer GPU, with bootstrap confidence intervals, Holm-corrected paired tests, and deployment-realistic variants of the key design choices. Three findings emerge. (i) Rubric-decomposed prompting beats holistic prompting for every model under batch min-max aggregation (though Qwen2.5-3B drops significantly on one prompt), and under mean aggregation two unrelated families land within 0.01 at the 1.5-1.7B scale. (ii) Mapping trait scores into the prompt range is fragile to grader calibration: one model compresses traits into a narrow low band (2-4 on 0-10) and naive mean aggregation collapses, while the min-max normalization of Multi-Trait Specialization repairs it (macro QWK 0.204 to 0.388) and stays within 0.03 when its statistics are frozen on 30 held-out essays. (iii) Signed error falls with essay length in eleven of twelve configurations, opposite to the verbosity bias reported for large LLM judges; normalized rubric decomposition largely flattens this slope for well-calibrated models. We anchor results honestly: the best local configuration (0.388) remains far below both the human inter-rater ceiling (0.769) and a length-only baseline (0.523), so we position sub-3B local models strictly for formative, human-supervised feedback.","authors":["Nguyen Dung Son","Dang Quang Minh","Nguyen Huu Loi","Truong Viet Vu","Nguyen Thai Anh"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20250","pdf_url":"https://arxiv.org/pdf/2609.20250","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["自动作文评分","小模型","零样本"],"reason":"论文是LLM自动作文评分，属于NLP能力评测，不以人类行为仿真为目标，无人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:39","error":null,"has_summary":false,"summary":null},{"id":"2609.19864","version":1,"title":"Converging Naming Styles, Persistent Network Locality: GitHub in the LLM Era","zh_title":"命名风格趋同，网络局部性持续：LLM时代的GitHub","abstract":"Social conventions often emerge through repeated interactions within social networks, allowing shared practices to coexist with variation across groups. Large language models (LLMs) introduce a potentially different coordination structure: a small number of widely used models can expose socially distant users to similar patterns and suggestions. Whether broad convergence under such shared technological influences eliminates network-local variation remains unclear. We examine this question in software development, where identifier naming styles provide observable conventions and LLM-based tools have rapidly diffused. Using public GitHub repositories created between 2015 and September 2025 across six programming languages, we characterize naming styles with 27 features and examine their association with detected LLM-related commits, their diversity across repository creation cohorts, and their relationship to owner proximity in a large-scale collaboration network. Repositories with detected LLM-related commits tend to use longer identifiers and, in several languages, make greater use of naming patterns already prevalent within the language. We also observe lower naming-style diversity in recent creation cohorts, with marked declines appearing around 2023-2024 in several languages, although their timing and trajectories differ. At the same time, network locality persists: in five of the six languages, repositories whose owners are closer in the collaboration network remain more similar in naming style even among recent, more homogeneous cohorts. These findings show that aggregate convergence and network-local variation can coexist, highlighting the need to examine not only how much cultural variation remains, but also how that variation continues to be structured by human social relationships in the era of widely shared AI systems.","authors":["Yuto Tamura","Sho Tsugawa"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19864","pdf_url":"https://arxiv.org/pdf/2609.19864","source_feed":"cs.SI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM影响","代码风格","社会网络"],"reason":"研究LLM对代码命名风格的影响，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2609.11286","version":2,"title":"Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data","zh_title":"生成一致的企业：多系统业务数据的合成与无参考评估","abstract":"Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model, a set of business applications, and a random seed, it produces a complete fictional enterprise: a workforce, a customer base, sales deals, support tickets, recorded calls, chat messages, and documents, all consistent with one another. One entity graph is projected into the native formats of 66 business products, so the same customer appears in the CRM, the support desk, and the call system under one identity. Because no real counterpart exists, realism is built in from cited reference statistics and verified by reference-free measurement: a five-axis scorecard of 28 statistical checks, an adversarial detector that hunts for the marks of synthetic generation, and a set of soundness checks that include a classifier test against an independently shuffled copy of the data. Because these instruments existed before the generator was tuned, progress is measured under a fixed yardstick: over 23 generated companies, mean realism climbed from 60.3 to 99.1, the weakest company from 41.1 to 94.9, and the detector, which initially flagged 55.2% of all records, now flags none. The scores hold on a seed never used during development. A second generator builds relational databases from a list of business questions. It forces qualifying rows for each answerable question, adds controlled near misses, and computes exact labels from the finished tables. The generator runs as a hosted service at https://console.era.eon.io. A company built there to a specification is served through its simulators over MCP and REST, and the simulators are also published as container images for offline use","authors":["Benjamin Gruenbaum","Doron Porat","Assaf Natanzon","Roy Zavida","Chen Dinachi","Or Itzahary","Omer Niv"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-12","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.11286","pdf_url":"https://arxiv.org/pdf/2609.11286","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["合成数据","企业数据生成","软件测试"],"reason":"生成企业合成数据用于软件测试，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:11","error":null,"has_summary":false,"summary":null},{"id":"2609.19989","version":1,"title":"Benchmarking LLM Compliance with China AI Generated Content Regulations","zh_title":"评估大语言模型对中国AI生成内容法规的合规性","abstract":"The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China's current AI-Generated content compliance requirements and provides evaluation results on 20 notable LLMs, offering insight into China's regulatory landscape. We design a novel framework to assess the compliance and refusal rates with 2303 questions spanning six distinct dimensions, including 203 self-constructed constitutional questions. The framework employs several judges to generate verdicts independently based on their hierarchical alignment memory. Our findings show that international models also exhibit high levels of compliance despite the use of standard Chinese questions, and the main differences may stem from dimensions closely related to ideological alignment. We establish a regulatory benchmark that enables the global AI community to evaluate both Chinese and non-Chinese LLMs under a unified set of legally grounded compliance requirements.","authors":["Chenrui Cui","Hongye Fang","Lisha Song","Weichao Chen","Yue Zhu","Gang Xu"],"categories":["cs.CL","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19989","pdf_url":"https://arxiv.org/pdf/2609.19989","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["合规性评测","内容监管","模型能力"],"reason":"评估LLM对中国AI内容法规的合规性，属于模型能力评测，不以人类行为为参照，不…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2609.20207","version":1,"title":"Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices","zh_title":"随机词汇演算基础：概率单纯形上的语义下降与随机动力学","abstract":"Large language models produce prompt-dependent probabilities over words, whereas scientific systems require uncertainty over meaningful states that can be updated as evidence arrives. We develop an observable framework for determining when language-derived probabilities support such a sequential state representation. Theoretically, we define typed measurable transformations of contextual language, construct a minimal closed representation, and give necessary and sufficient conditions for semantic updates to exist uniquely. We bound irreducible nonclosure and accumulated error, and under average contraction prove existence, uniqueness and stability of an external random recursion on a probability simplex. These results define a stochastic lexical calculus without attributing an internal calculus to the language model. Empirically, frozen experiments test the observable implications. Raw prompt-conditioned probabilities fail the prespecified invariance gate; after prompt-specific calibration, a common three-state representation passes the stability gates and covers 28 of 30 untouched eight-step paths, or 0.933 at nominal level 0.90. Accordingly, language probabilities support a stochastic state only conditionally on verified closure, stability and coverage within a declared operating domain.","authors":["Matthew F Dixon"],"categories":["cs.CL","math.CT","math.PR","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20207","pdf_url":"https://arxiv.org/pdf/2609.20207","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型概率","数学理论","语义表示"],"reason":"研究语言模型概率的数学性质，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:37","error":null,"has_summary":false,"summary":null},{"id":"2609.20584","version":1,"title":"SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment","zh_title":"SAFARI：面向LLM辅助危险分析与风险评估的工业基准","abstract":"Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial benchmark for LLM-assisted automotive Hazard Analysis and Risk Assessment (HARA) under ISO 26262. It contains 3,000 de-identified industrial HARA cases and evaluates two coupled tasks: open-ended hazard analysis and standards-grounded risk assessment. To evaluate open-ended HARA artifacts, we propose the first reference-anchored LLM-as-a-judge protocol with high expert correlation. Experiments with nine frontier LLMs show that models often produce plausible hazard narratives but remain weak at ISO 26262 risk classification, with the best ASIL macro-F1 reaching only 0.261. Chain-of-Thought prompting provides limited benefit and often degrades categorical risk assessment. Error analysis further localizes major failures to scenario-critical context omissions during hazard generation and to controllability misjudgments during risk assessment, indicating where expert oversight should be concentrated. The dataset can be obtained from https://github.com/xixi47520-hash/HARA.","authors":["Chenxi Wu","Zimu Wang","Haiyang Zhang","Wei Wang","Zhijie Xu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20584","pdf_url":"https://arxiv.org/pdf/2609.20584","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM辅助安全工程","自动驾驶功能安全","风险评估基准"],"reason":"论文聚焦汽车功能安全中的危险分析与风险评估，属于自动驾驶安全工程，不涉及用LL…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:42","error":null,"has_summary":false,"summary":null},{"id":"2609.20684","version":1,"title":"HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication","zh_title":"HerHealthEval：评估女性健康沟通的多语言与语域敏感理解","abstract":"Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.","authors":["Hassan Saeed Hassan Albattra","Mazen Mohammed Bahgat","Rahatara Ferdousi","Hana Essam Sayed Ahmed Amrya","Mariam Mousa"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20684","pdf_url":"https://arxiv.org/pdf/2609.20684","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","医疗健康","多语言"],"reason":"评估LLM对女性健康文本的理解，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2609.20808","version":1,"title":"Unifying Models of Intergroup Hostility in Online Discourse","zh_title":"统一网络话语中群体间敌意的模型","abstract":"Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. However, these theories were developed largely in parallel, often propose different and sometimes conflicting accounts of how hostility develops, and have rarely been tested against each other in real discourse. The result is a fragmented understanding of the rhetorical mechanisms of hostility, without a clear sense of how they appear, and relate to each other, in real-world discourse. Using 2.86 million posts from TikTok, Truth Social, and Twitter/X during the 2024 U.S. presidential election, we model the mechanisms of six foundational theories of intergroup hostility -- boundary construction, threat construction, scapegoating, negative evaluation, dehumanization, and action orientation -- within a common empirical framework to recover the broader organization of intergroup hostility rhetoric. Structurally, we find that boundary construction and threat construction anchor the system; temporally, we find that these mechanisms tend to follow a regular ordering: boundary construction, derogation, and action orientation tend to appear early; dehumanization and threat construction later; scapegoating latest. Mapping how these theoretical frameworks actually manifest in discourse bridges longstanding divisions across social science traditions and presents computational social science with a clearer empirical foundation for modeling intergroup hostility rhetoric beyond single-label detection.","authors":["Patrick Gerard","Julia Mendelsohn","Kristina Lerman"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20808","pdf_url":"https://arxiv.org/pdf/2609.20808","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["计算社会科学","群体间敌意","社交媒体分析"],"reason":"论文分析真实社交媒体帖子，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:46","error":null,"has_summary":false,"summary":null},{"id":"2609.20821","version":1,"title":"Embedding Models Measure in Peculiar Ways","zh_title":"嵌入模型以奇特方式度量","abstract":"Embedding spaces define notions of semantic similarity and distance. We study whether those embeddings reflect physical measurements of mass, distance, time and volume, which admit a unique, objective notion of semantic equivalence and distance. We find that physical measurement is only weakly modeled in the embedding space, and that instead quite peculiar measurement patterns can be observed. Further analysis indicates that embedding representations of physical measurements are strongly influenced by superficial string similarity, and recalibration of similarity does not substantially improve the alignment.","authors":["Juri Opitz","Andrianos Michail"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20821","pdf_url":"https://arxiv.org/pdf/2609.20821","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["嵌入空间","语义相似度","模型评估"],"reason":"研究嵌入空间对物理测量的表征，属NLP模型能力分析，不以人类行为为参照，与LL…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:47","error":null,"has_summary":false,"summary":null},{"id":"2609.19244","version":1,"title":"Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses","zh_title":"对话式LLM代理的网页搜索特征：从搜索决策与策略到结果与响应","abstract":"Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.","authors":["Mahsa Amani","Seungeon Lee","Abhisek Dash","Asmaa El Fraihi","Yunah Jang","Elisabeth Kirsten","Qinyuan Wu","Krishna P. Gummadi","Manish Gupta","Abhilasha Ravichander","Muhammad Bilal Zafar","Soumi Das"],"categories":["cs.AI","cs.IR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19244","pdf_url":"https://arxiv.org/pdf/2609.19244","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","网页搜索","对话系统"],"reason":"研究LLM代理的网页搜索行为，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.20001","version":1,"title":"E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews","zh_title":"E-AVI：基于证据的多模态自动视频面试评估","abstract":"Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.","authors":["Haoshen Wang","Dongbo Che","Zeyi Xie","Yuanjie Du","Shicheng Hua","Xingyu Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20001","pdf_url":"https://arxiv.org/pdf/2609.20001","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["自动面试评估","多模态学习","可解释AI"],"reason":"论文研究自动视频面试评估，属于多模态预测，不涉及用LLM仿真人类被试或与人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2609.20152","version":1,"title":"MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents","zh_title":"MTVA-Bench：评估级联语音代理中的语言模型","abstract":"Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller's voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.","authors":["Pritish Mishra","Ishaan Kumar","Akshat Mandoli","Sudarshan Kamath"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20152","pdf_url":"https://arxiv.org/pdf/2609.20152","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["语音代理","基准测试","多智能体系统"],"reason":"评估语音代理中的语言模型，属于多智能体系统评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2609.19831","version":1,"title":"Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles","zh_title":"复现透明可解释的推荐：通过自然语言用户画像探索开放权重模型","abstract":"In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preferences. The original paper explores the synthesis of user profiles from raw user-generated review text across domains such as movies and accommodations (Amazon Movies & TV, TripAdvisor). Crucially, these natural-language user profiles enable direct user interaction and intervention, allowing users to customize recommendations by correcting misattributed preferences or addressing cold-start settings. We successfully reproduce the core findings of the original study. Additionally, we extend the evaluation by conducting systematic context ablation experiments, multi-seed stability across five distinct random seeds to establish statistical reliability, and a mechanistic interpretability analysis using the nnsight framework to probe internal model representations under counterfactual profile perturbations. Our findings verify the original paper's claim that User Profile Recommendation (UPR) achieves competitive performance under its test-set reranking protocol and makes recommendations more transparent. Perturbing the natural-language profiles does change predictions, but it shifts predicted ratings uniformly across genres with no detectable genre-selective effect, leaving rankings unchanged even under direct activation steering. We trace this back to the rating-regression objective rather than the profile interface, with ranking-objective models clearly exceeding in this task.","authors":["Noah Mami\\'e","Laurin van den Bergh"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19831","pdf_url":"https://arxiv.org/pdf/2609.19831","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["推荐系统","可解释性","用户画像"],"reason":"论文研究推荐系统透明性，用LLM生成用户画像，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2609.20218","version":1,"title":"Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data","zh_title":"LLM时代训练经典模型还值得吗？表格数据上的交叉基准测试","abstract":"Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic's Claude for Excel - raising a practical question for the many business prediction problems where labels are expensive: should you prompt a frozen LLM, or collect data and train a model - and if so, how much data? We quantify the answer with the labeled-data crossover N*, the training-set size at which a trained classical model's learning curve overtakes a frozen LLM's training-free (and therefore flat) error. Aggregating 126 independent student evaluations of small GPT models under eight prompting configurations across 18 tabular datasets, paired with authoritative power-law learning curves for six classical model families, we find that training wins fast: even given an oracle choice of its best prompt configuration, a trained classical model beats the small frozen LLM using no more labeled data than is already on hand in 86% of cases, and wins by the smallest labeled subset we evaluate in 40%, with the observed crossover at a median of ~6% of the training set. In-context few-shot examples do not behave like training - error versus shot count does not follow a power law - and the same protocol re-run by independent implementers varies with a coefficient of variation of 0.148. A controlled probe indicates the LLM depends on recognizable feature-name semantics, which plausibly makes our crossover a conservative estimate (we do not claim memorization). For a typical business table, the evidence is clear: collect a few hundred labels and train a gradient-boosted model.","authors":["Kaihua Ding"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20218","pdf_url":"https://arxiv.org/pdf/2609.20218","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["表格数据","模型比较","基准测试"],"reason":"论文比较LLM与经典模型在表格数据上的预测性能，属于纯NLP能力评测，不以人类…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:38","error":null,"has_summary":false,"summary":null},{"id":"2609.20620","version":1,"title":"A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies","zh_title":"AUV故障恢复仿真平台：探索基于LLM的诊断策略","abstract":"Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.","authors":["Khalid Halba","Kylie Cooper","James G. Bellingham"],"categories":["cs.RO","cs.AI"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20620","pdf_url":"https://arxiv.org/pdf/2609.20620","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AUV","故障恢复","LLM诊断"],"reason":"论文研究AUV故障恢复，LLM用于诊断规划，属于机器人仿真环境，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2609.19364","version":1,"title":"Durably Reducing Belief in Women's Health Misinformation Through Culturally Adaptive AI Videos","zh_title":"通过文化适应性AI视频持久降低女性健康错误信息信念","abstract":"Health misinformation disproportionately harms women, yet interventions rarely address the community norms that sustain false beliefs. We test whether culturally adaptive AI-generated video in which the presenter looks like someone from her community reduces misinformation belief among low-literacy women in suburban India. In a field experiment (N=434), participants watched an AI-generated video featuring either an adaptive or neutral presenter. The culturally adaptive presenter reduced misinformation belief by 30%, nearly twice the reduction produced by the neutral presenter compared to the non-intervention control condition. Post-experiment interviews suggest women recalled the neutral condition as a generic video but recognized the adaptive presenter. Gains persisted for three weeks. The adaptive advantage was largest for beliefs reinforced by community, such as blaming women for infertility, and negligible for medical knowledge gaps, such as understanding vaccines. These findings demonstrate the potential of culturally adaptive AI interventions to counter socially embedded health misinformation.","authors":["Anku Rani","Kokil Jaidka","Shruti Sharma","Pragya Mahajan","Manisha Wadhwa","Andrew B. Lippman","Pattie Maes","Paul Pu Liang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19364","pdf_url":"https://arxiv.org/pdf/2609.19364","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["健康传播","AI视频","错误信息干预"],"reason":"研究AI视频干预健康错误信息，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2609.19420","version":1,"title":"Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026","zh_title":"大型语言模型在同行评审中的使用与效果：ICML 2026随机实验与调查","abstract":"LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.","authors":["Sunnie S. Y. Kim","Wesley Hanwen Deng","Jennifer Wortman Vaughan","Buxin Su","Weijie Su","Alekh Agarwal","Sharon Li","Martin Jaggi","Daniel G. Goldstein","Nihar B. Shah","Miroslav Dud\\'ik"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19420","pdf_url":"https://arxiv.org/pdf/2609.19420","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["同行评审","LLM使用政策","随机实验"],"reason":"研究人类审稿人使用LLM的行为，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2609.20490","version":1,"title":"TeamCAMS: An Open-Source Research Platform for Studying Human Behaviour in Human-AI Teams","zh_title":"TeamCAMS：研究人机团队中人类行为的开源研究平台","abstract":"In this article, we present TeamCAMS (Cabin Air Management System), a collaborative work environment for simulating human-AI (artificial intelligence) interaction for scientific research. The article outlines how several psychological theories guided the development of this multiple-task simulation. Modelling a process control environment, previous versions of TeamCAMS have already been used in empirical studies to address a wide range of research questions (e.g., comparing different forms of automation, evaluating impact of automation reliability, effects of stress on multiple-task performance). Outlining the technical possibilities offered by TeamCAMS, the article points out how its latest version offers researchers the possibility of addressing a set of new research questions including problems associated with teamwork (e.g., within-team conflict, distributed teamwork) and human-AI interaction. Finally, we will outline how this simulation environment can be enhanced further still to address research questions in new fields (e.g., automation of leadership). To promote transparency, reproducibility, and further development, TeamCAMS is made available to the research community under an open-source license.","authors":["Amos Brocco","Alain Chavaillaz","Andreas Sonderegger","Juergen Sauer"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20490","pdf_url":"https://arxiv.org/pdf/2609.20490","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","实验平台","团队协作"],"reason":"该平台用于人类与AI团队实验，非LLM仿真人类被试，且无LLM作为被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:41","error":null,"has_summary":false,"summary":null},{"id":"2609.20167","version":1,"title":"Utilizing AI-Driven Project Management Tools for Optimized Talent Management in HRM: A Framework for Enhanced Resource Allocation and Performance Prediction","zh_title":"利用AI驱动的项目管理工具优化人力资源管理中的人才管理：一个增强资源分配和绩效预测的框架","abstract":"When it comes to aiding businesses with demanding tasks regarding human resource management, TalentOptima unequivocally boasts of the best there is to offer. This tool utilizes AI based decision making, advanced predictive analytics, and also machine learning, all of which help in enabling automated resource allocation. To aid with better human resource management, TalentOptima integrates perfectly with already existing HR frameworks such as tools, etc. and shifts the focus towards aiding the user with insights while simultaneously alleviating manual work, this aids in a plethora of positive HR outcomes. A total of 40 managers participated in a simulation via user testing to ascertain if HR costs would reduce and work productivity would rise, the results were quite clear, attrition rates had dipped alongside risk and resource management rates, TalentOptima was a clear winner. Whereas the other HR frameworks primarily focused on ensuring work was done, TalentOptima ensured optimal and innovative decision-making, which overtime has proven to be invaluable for multiple companies, these results aid in proving why the tool is revolutionary.","authors":["Jay Barach"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20167","pdf_url":"https://arxiv.org/pdf/2609.20167","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI项目管理","人力资源管理","资源分配"],"reason":"论文是AI项目管理工具在HRM中的应用，不涉及LLM仿真人类被试，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:36","error":null,"has_summary":false,"summary":null},{"id":"2609.19225","version":1,"title":"Making Local Government Contracts Legible: A Computational Pipeline for Classifying and Mapping Intergovernmental Service Agreements","zh_title":"让地方政府合同可读：一种用于分类和映射政府间服务协议的计算流水线","abstract":"Interlocal agreements are one of the primary instruments through which local governments formalize collaboration for public service delivery, yet the institutional and financial content encoded in these contracts has remained inaccessible to systematic analysis at scale. This paper introduces an end-to-end computational pipeline for classifying intergovernmental agreements by institutional form and extracting financial relationships between principals and agents in service contracts. Applied to Iowa's 28E archive (N = 21,629), the largest dataset of interlocal agreements in the United States, the pipeline combines LLM-based summarization and classification across LLaMA 3.1, GPT 5.2 Pro, and Gemini 3 Pro on a four-class classification task that distinguishes agreements as either service contracts, resource sharing agreements, joint operations agreements, or new joint entity agreements. We also identify the financial principal and agent in these agreements and contracts, as well as the resulting dollar amounts and represent them on a directed network. The resulting financial network is organized around a small number of dominant service providers, with counties serving as the most structurally versatile actors, and cities as predominantly principals. By rendering the content of Iowa interlocal agreements analyzable at scale for the first time, this pipeline establishes a reusable methodology that researchers and state agencies can apply to track how public dollars move across local governments and to identify entities that depend heavily on a small number of providers.","authors":["Mohsen Ghasemizade","Ra\\'ul Guti\\'errez-Meave","Cailin Gramling","Aviral Chawla","Michael Robinette","Kate Albrecht","Juniper Lovato"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19225","pdf_url":"https://arxiv.org/pdf/2609.19225","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","合同分类","政府数据"],"reason":"使用LLM进行合同分类与信息提取，属于NLP应用，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.20211","version":1,"title":"Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines","zh_title":"沉默即认可：LLM智能体管道中的验证状态洗白","abstract":"Safety monitors in LLM agent systems often judge actions from summaries or stored handoffs, not from the original evidence. This creates a simple but dangerous failure mode: the handoff preserves the claim that an action is authorized while losing the fact that the claim was never verified. We call this verification-status laundering. Across nine open-weight monitors and two hosted models, the action and authorization proposition remain fixed while we remove the unverified provenance framing around the claim. This change raises approval for risky actions from $5\\%$ to $60\\%$ on Llama-3.1-8B and from $9\\%$ to $98\\%$ on Qwen2.5-14B, with similarly large shifts on both hosted models. The failure also emerges in ordinary agent pipelines. Summarizers frequently weaken the status, memory compressors often remove it, and a full proposer--summarizer--memory--monitor pipeline raises risky approval to $57$--$81\\%$ across three downstream monitors. Experiments on WildGuard and ATBench show the same pattern on independently authored harmful and unsafe requests: unsupported authorization claims make approval substantially more likely. Explicitly instructing monitors to reject unverified authorization is not a reliable cross-model fix: some models remain vulnerable, while others reject legitimate requests. Agent systems should therefore carry authorization provenance as structured state attached to the claim throughout the pipeline.","authors":["Yibo Hu"],"categories":["cs.CR","cs.MA"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20211","pdf_url":"https://arxiv.org/pdf/2609.20211","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全","多智能体系统","验证状态"],"reason":"研究LLM agent管道中的安全监控漏洞，属多智能体系统安全，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:37","error":null,"has_summary":false,"summary":null},{"id":"2609.20249","version":1,"title":"Accuracy Is Not Enough: A Cross-Architecture Audit of Demographic Bias in Deep Knowledge Tracing","zh_title":"准确性不足：深度知识追踪中人口统计偏差的跨架构审计","abstract":"Deep knowledge tracing (DKT) models implicitly decide which students an adaptive system believes have mastered a skill, yet almost all evidence on their demographic fairness comes from Bayesian knowledge tracing; the deep models that power modern systems have received no comparable cross-architecture audit. We close this gap: four architectures (DKT, DKVMN, SAKT, AKT) trained under three regimes (standard, reweighting, adversarial) on two public datasets with demographic metadata, Eedi (15.9M interactions) and OULAD (167k after preprocessing), evaluated with ABROCA, student-level bootstrap confidence intervals, and permutation tests addressing recent critiques of fairness-metric instability. Three findings emerge. (i) Bias is real but context-dependent: every architecture shows a significant socioeconomic ABROCA on Eedi (0.018-0.023, $p<0.005$), with per-group AUC lower for economically disadvantaged students, while gender bias is significant on OULAD for three of four architectures after multiplicity correction yet negligible on Eedi. (ii) The most accurate architecture is the most biased: AKT gains about 4 AUC points from item-level Rasch embeddings and shows the largest socioeconomic ABROCA, exceeding every other architecture under a paired bootstrap ($p\\leq0.002$); ablating only the Rasch embeddings removes the accuracy gain and the excess bias together. (iii) Standard mitigation is unreliable: reweighting and adversarial debiasing leave ABROCA essentially unchanged in every configuration that preserves accuracy, even though the adversary is pinned at chance at full reversal strength and a weak-strength positive control rules out a dead probe.","authors":["Dang Quang Minh","Nguyen Dung Son","Nguyen Huu Loi","Truong Viet Vu","Nguyen Thai Anh"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20249","pdf_url":"https://arxiv.org/pdf/2609.20249","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["知识追踪","公平性审计","教育数据挖掘"],"reason":"论文研究深度知识追踪模型的公平性审计，不涉及用LLM仿真人类被试，属于教育数据…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:38","error":null,"has_summary":false,"summary":null},{"id":"2609.20101","version":1,"title":"Competing for a Finite Pool of Attention in Social Media? How a New Geopolitical Conflict Reshapes Engagement in Bluesky","zh_title":"争夺社交媒体中的有限注意力？新地缘政治冲突如何重塑Bluesky中的参与","abstract":"Major geopolitical crises can rapidly reshape online public attention. Yet population-level increases in discussion volume about a new crisis reveal little about how users accommodate this new demand for attention. We study the onset of the Iran-US-Israel conflict, triggered on 28 February 2026, using longitudinal repost activity from Bluesky across four consecutive approximately three-month windows spanning the period before and after its onset; the data comprise 91.0 million unique posts and 645.5 million repost observations. We find that the new conflict reorganized participation through both reallocation among existing conflict participants and substantial activation of previously low-conflict-active users, while some previously active users reduced their conflict-related participation. Attention redistribution differed substantially across pre-existing interests: Iran-US-Israel and Israel-Palestine attention showed strong positive co-movement with little systematic relative replacement, whereas Other Political and Non-Political content more consistently lost attention share, and Russia-Ukraine exhibited weaker, heterogeneous replacement. Finally, disruption of users' broader attention allocation was substantially more prevalent among users with established attention to geopolitical conflicts than in the overall or Non-Political populations. Together, these findings show that a newly emerging conflict reorganizes online attention through turnover in who participates, selective co-attendance or replacement across topics, and disruption of broader attention patterns concentrated among users already engaged with geopolitical conflicts.","authors":["Kamand Kalashi","Arash Badie-Modiri","Ali Salloum","Juhi Kulshrestha","Talayeh Aledavood","Mikko Kivel\\\"a"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20101","pdf_url":"https://arxiv.org/pdf/2609.20101","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社交媒体分析","注意力分配","地缘政治冲突"],"reason":"研究社交媒体注意力再分配，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2609.20655","version":1,"title":"Using machine learning metrics to provide deeper insights into the performance of choice models","zh_title":"使用机器学习指标深入洞察选择模型性能","abstract":"Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance.","authors":["Lorenzo Mu\\~noz","Stephane Hess","Thomas O. Hancock","Georges Sfeir"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20655","pdf_url":"https://arxiv.org/pdf/2609.20655","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["选择建模","机器学习","模型评估"],"reason":"论文讨论选择模型的评估指标，不涉及LLM仿真人类被试，属于纯方法论研究。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2606.27845","version":2,"title":"LLM Agents as Static Level-k Players in Behavioural Games","zh_title":"行为博弈中作为静态层级-k玩家的LLM智能体","abstract":"Large Language Models (LLMs) are increasingly used as stand-ins in behavioural games. These stand-ins rely on the assumption that the LLM's distribution of choices meaningfully matches how humans play the same game. This study tests that assumption through two games. The first is a p-beauty contest, and the second one is a public goods game. The study first investigates five local-model settings within the same model family. These settings are varied together in a 360-cell factorial, which balances temperature, scale (0.5-32B), quantisation, instruct vs base, and framing. Each cell's distribution is then compared against whole choice distributions in published human data. Each deployment setting, except for quantisation, governs a different aspect of fidelity. Mechanically, while the dispersion of human players can be somewhat recovered through deployment settings, the strategic process behind it cannot. Through the lens of the level-k cognitive theory, we find that LLMs act as static, category-retrieved level-k players, where k is set by the model scale. The models also do not run within-game belief-updating or backward induction throughout multiple-round horizon settings. While human contributions decayed in the public goods game, LLMs stayed flat or rose at every scale. When the horizon test was administered, LLMs were more cooperative under an indefinite horizon compared to a finite one. However, LLMs ignore their relative round position, so no last-round defection was displayed. This implies that LLMs retrieved levels relative to the horizon category rather than working out iteratively from the specific game setting.","authors":["Po Han Teo"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-17","first_seen":"2026-06-26","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2606.27845","pdf_url":"https://arxiv.org/pdf/2606.27845","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"直接测试LLM在行为博弈中替代人类被试的保真度，并与已发表人类数据对照，发现静…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":1,"question":"LLM在行为博弈中作为人类被试替代品时，其选择分布和策略过程是否与真实人类一致？","design":"使用Qwen 2.5模型家族，在p-beauty contest和公共品博弈中，通过360个单元格的因子设计操纵温度、模型规模（0.5-32B）、量化、指令/基础模型和框架，测量LLM的选择分布，并与已发表的人类数据比较。","baseline":"已发表的人类行为数据，包括p-beauty contest和公共品博弈中的选择分布。","findings":"LLM表现为静态的、类别检索的level-k玩家，k由模型规模决定；LLM没有进行回合内信念更新或逆向归纳，在公共品博弈中贡献不衰减，且忽视相对回合位置，无最后一轮背叛。","reliability":"论文未讨论","relevance":"该研究直接检验LLM在行为博弈中替代人类被试的保真度，并与真实人类数据对照，发现LLM的策略过程与人类不同，对评估LLM仿真可靠性至关重要，值得精读原文。","inspiration":"采用因子设计系统操纵LLM部署设置，并与人类分布整体比较的方法值得借鉴｜可迁移到政策公告的预期形成实验，如央行沟通对通胀预期的影响｜用LLM模拟公众对政策公告的反应，处理为不同政策措辞或沟通方式，结果变量为预期通胀分布，与调查预期数据（如密歇根大学调查）对照。"}},{"id":"2609.17549","version":1,"title":"Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues","zh_title":"社会模式在合成数据中是否成立？分析LLM生成与真实对话中的网络欺凌动态","abstract":"Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether such data faithfully reproduces the social dynamics of authentic interactions beyond supporting downstream task performance. We present a comprehensive framework for evaluating the social realism of LLM-generated CB conversations. We compare authentic and synthetic dialogues generated by GPT, Grok, and LLaMA across interactional structure (turn-taking, power dynamics, and repair behavior), linguistic and stylistic realism (pronoun usage and humor), affective and behavioral markers (CB types, profanity, and toxicity), and temporal escalation dynamics. We further complement automatic analyses with a human evaluation of cyberbullying presence, scenario relevance, role plausibility, and social realism. Our results show that LLM-generated data consistently preserves high-level interactional structure, including role participation patterns, directional power asymmetry, and broad distributions of behavioral markers. However, all models systematically distort finer-grained social phenomena, including behavioral magnitude, role-specific allocation, categorical distributions, and temporal dynamics. These distortions are strongly model-dependent: GPT suppresses harmful content, Grok amplifies aggressive behaviors, and LLaMA provides the most balanced approximation while smoothing role distinctions. Our findings show that synthetic CB data is useful for modeling global interactional structure but remains an imperfect substitute for authentic conversations when behavioral realism and social dynamics are essential.","authors":["Arefeh Kazemi","Hamza Qadeer","Sinan Asci","Joachim Wagner","Brian Davis"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17549","pdf_url":"https://arxiv.org/pdf/2609.17549","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会动态","真实性评估"],"reason":"用LLM生成对话仿真网络欺凌社会动态，并与真实对话对照，评估仿真保真度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":2,"question":"LLM生成的网络欺凌对话是否忠实再现了真实对话中的社会动态，而不仅仅是支持下游任务性能？","design":"使用GPT、Grok和LLaMA生成合成网络欺凌对话，与真实对话对比，分析互动结构（话轮转换、权力动态、修复行为）、语言风格（代词使用、幽默）、情感行为标记（欺凌类型、脏话、毒性）和时间升级动态，并进行人类评估。","baseline":"真实网络欺凌对话数据集（具体名称未在节选中提及）。","findings":"LLM生成的数据保留了高层互动结构，如角色参与模式、方向性权力不对称和行为标记的广泛分布；但所有模型都系统性地扭曲了细粒度社会现象，包括行为强度、角色特定分配、类别分布和时间动态，且扭曲程度因模型而异：GPT抑制有害内容，Grok放大攻击行为，LLaMA提供最平衡的近似但平滑了角色差异。","reliability":"论文承认合成数据在需要行为真实性和社会动态时是真实对话的不完美替代品，且扭曲具有模型依赖性。","relevance":"该研究直接评估LLM仿真社会互动的保真度，与研究者关注的人类仿真实验和可靠性评估高度相关，提供了系统的对照基准和偏差分析，值得精读。","inspiration":"借鉴其多维评估框架，将自动分析与人类评估结合，系统比较合成与真实数据在结构、语言、行为和时间动态上的差异｜可迁移到经济金融中的社会互动场景，如谈判博弈、团队决策或市场情绪传播｜设计一个实验：用LLM生成模拟投资者在社交媒体上的互动对话，处理为不同模型（如GPT、Claude）或提示策略，结果变量为情绪传染、羊群行为或信息扩散模式，与真实投资者论坛数据（如StockTwits）对照，评估仿真保真度。"}},{"id":"2609.18106","version":1,"title":"Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment","zh_title":"开放权重大语言模型应用于招聘时性别与种族偏见的语言触发因素","abstract":"Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.","authors":["Kosuke Kitahara","Nobuhiro Yamaguchi"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18106","pdf_url":"https://arxiv.org/pdf/2609.18106","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","招聘偏见","算法审计"],"reason":"用LLM仿真招聘中的人类决策，并与真实人类数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":5,"question":"招聘广告中的语言特征（如代理性/社群性词汇、编码排斥语言）如何触发开放权重LLM在招聘模拟中的性别与种族偏见？","design":"使用六个开放权重LLM（Llama 3.2、Mistral、Gemma 3、Qwen 3、Phi 3、DeepSeek-R1）进行招聘者模拟和求职者模拟实验。通过操纵招聘广告语言（代理性 vs. 社群性词汇、编码排斥语言）和候选人/求职者的人口统计标签（性别、种族），测量LLM给出的推荐评分、兴趣表达等结果变量，并进行标签消融实验和词嵌入关联测试。","baseline":"无对照（论文未使用真实人类数据作为基准，而是基于LLM输出进行审计）。","findings":"代理性招聘语言降低女性候选人的推荐评分，而社群性语言部分逆转该惩罚；编码排斥语言大幅抑制非白人候选人的推荐评分，并在求职者模拟中抑制非白人角色的兴趣表达，形成寒蝉效应。标签消融实验表明显式人口统计标签是主要因果驱动因素，词嵌入关联测试在表征层面证实了偏见。","reliability":"论文未讨论","relevance":"该研究直接针对LLM在招聘场景中的仿真行为，系统操纵语言变量并测量偏见输出，与您关注的LLM仿真可靠性及偏差评估高度相关，值得阅读原文以了解其审计协议和效应量。","inspiration":"借鉴其将文本特征作为处理变量、通过多模型审计和标签消融识别因果机制的方法。｜可迁移到信贷审批中的语言歧视研究，如贷款广告或申请表中的措辞对AI审批决策的影响。｜以LLM作为信贷审批员，处理为贷款申请描述中的代理性/社群性词汇或编码排斥语言，结果变量为审批通过率或利率，对照真实信贷审批数据（如抵押贷款披露数据）评估仿真偏差。"}},{"id":"2609.17933","version":1,"title":"AI Mediators Regulate Emotion and Create Value in Disputes","zh_title":"AI调解员在纠纷中调节情绪并创造价值","abstract":"In conflict and disputes, especially, emotion acts as a salient force in influencing outcomes. Prior work shows negative affect can obstruct collaborative behaviors, which typically lead to ``win-win'' outcomes. Thus, some suggest mediators may help regulate emotion and achieve joint gains. With the proliferation of AI, we posit LLMs may perform well at this task, with the added benefit of better accessibility compared with a human mediator. To examine the effectiveness of AI versus novice human mediators, we conduct a between-subjects experiment, where participants engage in a dispute mediated by a human, AI, or no mediator. We first analyze how well the mediators regulate emotions within a dispute -- finding AI mediators perform significantly better than humans at reducing negative emotion. We next examine whether AI mediators facilitate disputants better realizing joint gains in disputes with high integrative potential (IP) -- we find a marginally significant interaction between IP and condition (AI versus human), indicating LLMs may outperform humans at aiding disputants realize joint gains. Lastly, we perform an analysis of the messages the mediators sent, finding the AI sent significantly more messages suggesting trade-offs compared to the humans.","authors":["James Hale","Jonathan Gratch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17933","pdf_url":"https://arxiv.org/pdf/2609.17933","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","调解实验","人机对照"],"reason":"用LLM作为调解人替代人类，与真实人类被试互动并对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":3,"question":"在情绪化的纠纷调解中，AI调解员相比新手人类调解员能否更有效地调节情绪并帮助双方实现共同收益？","design":"使用大语言模型（LLM）扮演AI调解员，与新手人类调解员和无调解员条件进行被试间实验。参与者扮演买卖双方进行纠纷谈判，调解员在每轮发言后决定是否干预并发送消息。结果变量包括情绪调节效果（负面情绪减少）、共同收益实现程度（积分潜力IP与条件的交互）以及调解员消息内容（如提出权衡建议的频率）。","baseline":"新手人类调解员作为对照，以及无调解员条件作为基线。参与者为真实人类被试，其偏好通过分配100点来测量，并根据达成目标获得金钱奖励。","findings":"AI调解员在减少负面情绪方面显著优于人类调解员；在具有高整合潜力的纠纷中，AI调解员可能比人类调解员更能帮助双方实现共同收益（边际显著交互作用）。此外，AI调解员发送了更多建议权衡的消息。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类调解员，与真实人类被试互动并对照人类调解员表现，属于人类仿真实验，且涉及情绪调节和谈判结果，与研究者关注的经济学实验和政策评估场景高度相关，值得阅读原文了解实验细节和局限性。","inspiration":"借鉴其使用LLM作为干预代理与人类被试互动的实验设计，以及通过偏好分配和积分潜力测量共同收益的方法。｜可迁移到经济金融中的谈判或冲突解决场景，如劳资纠纷调解、商业合同谈判、消费者投诉处理等。｜设计一个实验：以LLM作为调解员，人类被试扮演谈判双方（如买方和卖方），处理为AI调解员、人类调解员或无调解员，结果变量为谈判达成协议的质量（如联合收益、满意度）和情绪变化，对照真实人类调解员数据，并利用已有谈判实验数据集（如KODIS）进行基准比较。"}},{"id":"2609.18060","version":1,"title":"AI Peers Exert Social Influence on Human Dishonesty in Groups","zh_title":"AI同伴对群体中人类不诚实行为施加社会影响","abstract":"Human dishonesty in group settings is highly susceptible to peer influence, particularly when incentivized. Although artificial intelligence (AI) evolves from passive tools into active collaborators, its impact on human moral behavior within groups remains underexplored. We addressed this gap through a two-phase randomized behavioral study (N=280 and N=360). We found AI agents exert substantial social influence comparable in magnitude to that of human peers. Specifically, participants reported more dishonestly when exposed to dishonest rather than honest normative cues. This effect is evident across injunctive, subjective, and descriptive social norms. Interestingly, the only significant adjacent behavioral change occurred when dishonest peer behavior first appeared, whereas further increases from one to four dishonest peers produced weaker and non-monotonic changes. Furthermore, participants rapidly converge on decision-making, showing modest increases in dishonest reporting through repeated exposure. These findings highlight the importance of managing the behaviors and normative signals communicated by AI group members.","authors":["Shuning Zhang","Xinyuan Zhou","Yuanyang Qiu","Tianqi Song","Yuting Yang","Yiwen Ren","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18060","pdf_url":"https://arxiv.org/pdf/2609.18060","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","行为实验"],"reason":"用AI代理替代人类同伴，研究其对人类不诚实行为的社会影响，并与人类同伴对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":4,"question":"AI同伴的诚实规范如何影响群体中人类的不诚实行为？","design":"采用两阶段随机行为实验（N=280和N=360），使用激励性掷骰子报告任务。AI代理作为群体成员，通过描述性、指令性和主观性社会规范传递诚实或不诚实线索，测量人类被试的虚报行为。","baseline":"与人类同伴的社会影响进行对照，比较AI同伴与人类同伴的影响大小。","findings":"AI代理对人类不诚实行为产生显著社会影响，其影响程度与人类同伴相当。当出现不诚实同伴时，被试虚报增加，但同伴数量从1个增加到4个时影响减弱且非单调。","reliability":"论文未讨论","relevance":"该研究直接使用AI代理替代人类同伴，研究其对人类道德行为的影响，并与人类同伴对照，符合研究者对LLM仿真实验和真实人类基准的兴趣。","inspiration":"值得借鉴的是其将AI代理作为群体成员，通过操纵其行为规范来研究社会影响，并设置人类同伴对照组以量化AI影响的大小。｜可迁移到经济金融中的群体决策场景，如投资团队中AI顾问的不诚实建议对个人投资决策的影响。｜设计一个实验：被试与AI代理组成投资小组，AI代理提供虚报收益的建议（处理），测量被试的投资报告虚报程度（结果变量），并与人类同伴提供同样建议的对照组比较，同时使用真实投资数据作为外部基准。"}},{"id":"2609.07478","version":2,"title":"The Internal Anatomy of Strategic Choice in Large Language Models","zh_title":"大语言模型策略选择的内部解剖","abstract":"Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.","authors":["Vin\\'icius Ferraz","Leon Houf","Enrico Ferrea"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-09","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.07478","pdf_url":"https://arxiv.org/pdf/2609.07478","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","策略博弈","算法保真度"],"reason":"用LLM复现人类策略选择并与人类数据对照，分析内部计算与行为差异，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":6,"question":"大语言模型在一次性策略博弈中，内部表征与计算如何将激励信号转化为选择行为？","design":"用四个开源大语言模型（Qwen2.5 base/instruct、Llama-3.1-Instruct、GPT-OSS）作为被试，在144个严格序数2x2博弈中做一次性选择，记录激活值，用线性探针解码激励和选择，并进行激励强化干预和固定决策线索提示。","baseline":"人类选择数据来自Moore et al. (2026)的相同博弈和支付尺度，以及Zhu et al. (2025a)的基数支付数据集映射。","findings":"密集模型复现了人类随博弈复杂度下降的未调整选择模式；所有模型都能解码激励和选择，但激励是否到达选择、与选择对齐及干预效果因模型而异，基础与指令微调模型行为相似但内部路径不同。","reliability":"论文指出相似行为可能基于不同计算，后训练可重塑激励到决策的路径而保持行为和可解码信息基本不变，暗示行为仿真可能掩盖内部差异。","relevance":"该研究直接以LLM仿真人类策略选择并与真实人类数据对照，同时揭示内部计算与行为的不一致，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"借鉴其用线性探针追踪激励信号从输入到选择的内部路径，并施加干预检验因果性的方法，可迁移到经济决策仿真中检验模型是否真正使用经济激励而非表面模仿。｜可应用于资产定价实验，检验LLM是否内部表征风险溢价并据此决策。｜以LLM为被试，呈现不同风险收益的资产选择任务，用探针解码风险溢价信号并强化干预，与人类实验数据（如股票市场参与决策）对照，观察选择变化。"}},{"id":"2609.17534","version":1,"title":"Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits","zh_title":"LLM中的装好与装坏：黑暗三人格特质下的反应失真","abstract":"Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions. Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels. Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity. Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone. The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.","authors":["Victoria Popa","Guglielmo Cola","Caterina Senette","Maurizio Tesconi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17534","pdf_url":"https://arxiv.org/pdf/2609.17534","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格测量","反应偏差","仿真可靠性"],"reason":"研究LLM在人格测量中的反应失真，与人类数据对照，评估仿真偏差，可迁移到人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":7,"question":"LLM 在人格测量中是否会在假好（fake-good）和假坏（fake-bad）条件下系统性地调节黑暗三联征特质的表达？","design":"使用七个先进LLM，在就业选拔和司法评估两种情境下，通过提示框架施加假好或假坏动机，用TRAIT框架测量黑暗三联征（马基雅维利主义、自恋、精神病态）得分，并与基线自我评估比较。","baseline":"无对照","findings":"大多数模型在假好条件下降低黑暗三联征得分，在假坏条件下提高得分，但效应大小和一致性因特质和模型而异。马基雅维利主义和自恋的偏移最强且最一致，精神病态异质性更大；就业情境比司法情境产生更大效应，显式假坏指令比情境框架产生更强扭曲。","reliability":"论文未讨论","relevance":"该研究展示了LLM对情境动机的敏感性，可用于评估仿真中的社会期望偏差和印象管理，对使用LLM模拟人类被试时的可靠性有警示意义。","inspiration":"借鉴其通过情境框架施加动机处理并测量特质偏移的方法，可迁移到经济金融中的社会期望偏差场景（如信贷审批中的歧视、消费者道德行为）。｜可设计实验：用LLM扮演贷款申请人，在强调社会责任或利润最大化的不同银行政策下，测量其自我报告的诚信或风险偏好，并与实际信贷数据中的偏差对照。"}},{"id":"2609.18282","version":1,"title":"Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement","zh_title":"好得难以置信？诊断并缩小AI偏好与真实用户参与度之间的差距","abstract":"Large language models are increasingly used to generate and evaluate online content, yet it remains unclear whether the qualities they associate with higher engagement match what real users respond to. We study this question using 1.17 million answers to 25,978 questions from Zhihu, Quora, and Reddit, comparing real platform answers and AI-generated answers across four within-question engagement levels. We introduce Ontological Preference Measurement, which represents answers along three dimensions: logic, affect, and expression. We find a systematic gap between AI preference and real user engagement: as target engagement increases, LLMs add more explicit logical structure, while real user engagement is more strongly associated with affective and expressive salience. We call this tendency logic overbinding. Based on this diagnosis, we propose Ontology-Masked Reasoning Autoencoding (OMRA), a controlled intervention that masks and reconstructs over-explained spans while preserving stance, factual content, and coherence. Across four LLM families, OMRA reduces the measured gap by an average of 54.4%. In human evaluation, OMRA wins 62.4% of pairwise preference judgments against matched real platform answers, even though the real answers are more often judged to be human-written.","authors":["Xinglang Zhang","Yuanmeng Xiang","Yunyao Zhang","Zeliang Chen","Junqing Yu","Zikai Song"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18282","pdf_url":"https://arxiv.org/pdf/2609.18282","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","内容生成"],"reason":"用LLM生成内容并与真实用户互动数据对照，诊断AI偏好与人类行为差距，属于仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":9,"question":"LLM 在生成高互动内容时偏好的文本特征是否与真实用户互动行为一致？","design":"使用 LLM 生成针对四个互动等级的答案，与真实平台答案对比，通过本体偏好测量框架从逻辑、情感、表达三个维度量化文本特征，并施加 OMRA 干预以缩小差距。","baseline":"来自知乎、Quora 和 Reddit 的 117 万条真实答案，按问题内投票排名分为四个互动等级。","findings":"LLM 在追求高互动时过度增加显性逻辑结构，而真实高互动答案更依赖情感和表达显著性，作者称之为“逻辑过度绑定”。OMRA 干预平均缩小 54.4% 的差距，并在人类评估中胜过真实答案。","reliability":"论文未讨论","relevance":"该研究直接对比 LLM 生成内容与真实用户行为，诊断 AI 偏好与人类反应的系统性偏差，并尝试通过干预修正，对评估 LLM 仿真人类行为的可靠性具有参考价值。","inspiration":"借鉴其本体驱动的多维测量和干预设计，可迁移到经济金融领域的文本生成场景，如政策沟通、市场评论或金融建议。｜例如，研究 LLM 生成的央行政策声明或分析师报告是否与真实市场反应匹配。｜以 LLM 为被试，要求其生成不同目标市场反应的金融文本，测量逻辑、情感、表达特征，并与真实市场数据（如股价波动、交易量）对照，检验 AI 偏好与真实投资者反应的差距。"}},{"id":"2609.17989","version":1,"title":"Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders","zh_title":"AI代理为谁工作？角色分配引发LLM推荐中的赞助偏差","abstract":"Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise consumers who rely on their judgment, yet are deployed by platforms that benefit when sponsored listings are chosen. Sponsorship disclosures, designed to allow consumers to penalize paid placements, now reach the AI agent rather than the consumer, and the agent's evaluation of them is hidden from the consumer. Drawing on the fiduciary concept of conflict of duty, we argue that an agent's evaluation of a sponsored listing should not depend on which party deployed it. In controlled choice experiments, we manipulate assigned roles in the system prompt to name either a traveler or a booking platform as the agent's principal. Platform delegation significantly attenuates the penalty that agents apply to sponsored listings and weakens the skepticism that disclosure triggers in their reasoning traces. We replicate out findings across LLMs and reasoning depths. A second study decomposes the disclosure label and shows that the divergence between the two delegates widens significantly when the paid placement is attributed to the platform. Stricter terminology (\"Sponsored\" instead of \"Promoted\") lowers choice of paid listings but does not close this gap when the platform is named. The findings show that disclosure mandates designed for human consumers cannot by themselves protect consumers in AI-mediated commerce.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17989","pdf_url":"https://arxiv.org/pdf/2609.17989","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者决策","赞助偏差"],"reason":"用LLM模拟消费者决策，与人类对照，揭示角色分配导致的赞助偏差，属经济学实验场…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":8,"question":"AI代理在同时服务消费者和广告平台时，其推荐行为是否会因委托方身份（消费者vs平台）而出现赞助偏差？","design":"用LLM（Gemini 3.1 Pro等）模拟购物助手，在系统提示中操纵委托方身份（旅行者vs预订平台），呈现带赞助标签的酒店列表，测量选择赞助列表的概率及推理痕迹中的怀疑程度。","baseline":"无对照","findings":"平台委托显著减弱了LLM对赞助列表的惩罚，选择赞助列表的概率从消费者委托时的50.2个百分点降至29.2个百分点；当赞助标签明确归属平台时，平台委托的代理甚至表现出对赞助列表的强烈偏好（选择率74.2% vs 消费者委托的34.6%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者决策，揭示了角色分配导致的赞助偏差，属于经济学实验场景，与研究者关注的人类仿真实验高度相关，值得精读原文。","inspiration":"值得借鉴的是通过最小化系统提示中的角色分配来诱发LLM行为偏差，并分析推理痕迹以揭示机制｜可迁移到信贷审批歧视或金融产品推荐场景，检验AI代理在银行与借款人之间的利益冲突｜设计：用LLM扮演贷款顾问，系统提示中分别指定委托方为银行或借款人，呈现带“银行推荐”标签的贷款产品，测量选择概率，并与人类贷款顾问的真实选择数据对照。"}},{"id":"2609.17544","version":1,"title":"Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation","zh_title":"大语言模型与中医医师的对比：真实世界临床病例评估","abstract":"Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.","authors":["Jiacheng Xie","Xiaoting Tang","Yang Yu","Jinpu Li","Shouli Li","Congcong Jing","Yantao Yang","Zhiyong Zhao","Ziyang Zhang","Qilin Song","Guanghui An","Dong Xu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17544","pdf_url":"https://arxiv.org/pdf/2609.17544","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医疗决策","人类对照"],"reason":"用LLM替代医生进行临床决策评估，并与真实医生对照，属于人类仿真，但场景为医疗…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":10,"question":"在真实中医门诊场景中，大语言模型能否替代或辅助中医师进行辨证论治，其诊断与处方质量相比执业中医师如何？","design":"用16个大语言模型扮演中医师，对60个真实门诊病例生成诊断与处方报告；同时60名执业中医师对相同病例生成报告；所有报告匿名后由5名资深中医专家在9个诊断与治疗维度上评分，并分析处方模式、幻觉与安全性。","baseline":"60名执业中医师对相同60个病例生成的诊断与处方报告，经同一批专家按相同维度评分。","findings":"先进通用大模型在专家评分上总体高于医师对照组，尤其在医疗建议、治疗原则和部分诊断任务上表现更好；但处方层面在药物选择、剂量和治疗策略上存在差异，且定性安全审查发现幻觉和模板化输出。","reliability":"论文承认需要医生监督、安全约束和前瞻性临床评估；指出LLM可能产生幻觉、不安全的草药组合或不适当剂量，且当前评估基于回顾性病例，缺乏真实临床结局验证。","relevance":"该研究用LLM替代医生进行临床决策并与真实医生对照，属于人类仿真在医疗场景的应用，提供了真实世界数据与专家盲评的严格设计，值得阅读以了解仿真在专业决策中的有效性与偏差。","inspiration":"借鉴其匿名化处理与专家盲评设计，可确保仿真输出与人类输出在相同条件下比较，减少评估偏差｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请做出审批决策，比较其与真人信贷员的差异｜以LLM作为被试，处理为不同申请人特征（如种族、性别），结果变量为审批决定与理由，对照真实信贷审批数据或信贷员决策记录，检验LLM是否复现或放大人类偏见。"}},{"id":"2609.17550","version":1,"title":"No Usable Linear \"Capitulation Direction\" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback","zh_title":"两个小型LLM中不存在可用的线性“屈服方向”：激活引导声明的验证协议，以及跨家族对反驳下谄媚行为的研究","abstract":"Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and Llama-3.2-1B, over TriviaQA: the model answers, is challenged with one of four scripted pushback styles, and answers again. Conditioned on an initially correct answer, the models flip to a wrong answer in 41.8% and 43.1% of episodes. Which pressure works is a property of the model, not the pressure: the same within-question paired comparison (bare doubt vs. emotional appeal), specified in advance, is Bonferroni-significant in opposite directions across families (Qwen: bare doubt > emotional, OR 2.5, p=.040; Llama: emotional > bare doubt, OR 4.0, p=.001). Failure mode is also model-dependent: Llama abandons answers without recommitting at six times Qwen's rate (8.2% vs. 1.4%). Identical pushback repairs initially wrong answers only ~13% of the time; pushback is net epistemically destructive. We then ask whether capitulation is linearly decodable from the pre-response residual stream, a prerequisite for steering-vector interventions at that locus. A naive difference-in-means probe appears to succeed (in-sample AUROC 0.81/0.71), but a validation protocol combining question-level cross-validation, shuffled-label nulls, and a known-direction positive control shows the signal is overfitting: the best cross-validated AUROC is 0.582 in Qwen and 0.548 in Llama, both near or below their permutation thresholds and far under a pre-registered usability bar of 0.70, while the identical pipeline recovers a pushback-presence control direction at AUROC 1.000 in both. We further quantify a measurement hazard: substring grading underestimates capitulation by 18-24 percentage points. Code, prompts, transcripts, and analysis are released.","authors":["Saad Aamir","Muhammad Awais Bin Adil"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17550","pdf_url":"https://arxiv.org/pdf/2609.17550","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM行为","可靠性评估","激活引导"],"reason":"研究LLM在用户反驳下的行为变化，评估其可靠性，与人类仿真中的偏差问题相关。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":11,"question":"在用户反驳下，小型指令微调语言模型放弃正确答案（屈从）的行为模式是什么？屈从是否可由残差流中的线性方向解码？","design":"使用两个不同家族的小型指令微调模型（Qwen2.5-1.5B 和 Llama-3.2-1B），在 TriviaQA 数据集上进行多轮问答：模型先回答，然后受到四种预设反驳风格之一的挑战，再回答。测量结果变量包括：初始正确时翻转为错误的概率、初始错误时修复的概率、放弃不重新承诺的概率，以及屈从的线性可解码性（AUROC）。","baseline":"无对照","findings":"在初始正确的情况下，两个模型分别有41.8%和43.1%的回合翻转为错误答案；哪种压力有效是模型特有的，同一比较在家族间方向相反。线性探测看似成功，但经过交叉验证和置换检验后，屈从方向不可用（AUROC接近机会水平），而阳性对照方向可完美恢复。","reliability":"论文承认的局限包括：仅两个家族且规模小（1-1.5B）；单一数据集（TriviaQA）和单一语言（英语）；跨模型比较非问题配对；模板固定；LLM 裁判未经完整人工验证；Llama 的关键比较是在观察临时标签后指定的。","relevance":"该研究通过行为实验和严格的验证协议，揭示了 LLM 在用户压力下的屈从行为具有模型特异性，且线性方向不可靠，这对使用 LLM 模拟人类决策时的偏差评估和可靠性判断有直接参考价值，值得阅读原文。","inspiration":"借鉴其多轮交互设计、预设处理（反驳风格）、结果分类（翻转/修复/放弃）以及严格的验证协议（交叉验证、置换检验、阳性对照）来评估模型行为的稳健性。｜可迁移到经济金融中的政策沟通或建议采纳场景，例如 AI 财务顾问在用户质疑下是否改变投资建议，或消费者对 AI 推荐产品的信任变化。｜以 LLM 作为被试，模拟投资者在 AI 投资建议受到用户质疑时的反应，处理为不同风格的反驳（如权威质疑、情感诉求），结果变量为是否改变初始投资决策，并与人类投资者在类似情境下的真实行为数据（如实验或调查数据）进行对照。"}},{"id":"2609.18068","version":1,"title":"From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale","zh_title":"从基列河到大型语言模型的推理分布：隐性方言偏见与大规模语言画像","abstract":"Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations in internal probability distributions untouched. Adapting the matched-guise sociolinguistic paradigm, we examine covert dialect bias in housing-related social judgments across four varieties: Standard American English (SAE), African American Vernacular English (AAVE), Nigerian Standard English (NSE), and Nigerian Pidgin (NP). AAVE reflects the racialized dialect studied in prior covert-bias evaluations, whereas NSE and NP represent Black African, postcolonial varieties absent from this literature. Using 260 meaning-matched sentence quadruples and log-probability scoring over housing-relevant adjectives, we probe ten open-weight LLMs across three contexts varying in social proximity: tenant screening, neighbor acceptance, and roommate selection. Across all ten models, AAVE and NP are consistently associated with more negative adjectives than SAE, with NP penalized most severely. Crucially, each dialect is penalized via distinct stereotype clusters rather than a generic non-standard category. NSE, which carries institutional prestige, displays a context-dependent shift: favored over SAE in formal tenant screening but increasingly penalized as social proximity grows. Our findings reveal that LLMs inherit covert dialect bias along both racial identity and prestige dimensions, echoing documented human housing discrimination and demonstrating its reach across postcolonial English varieties.","authors":["Chowdhury Mohammad Abdullah","Rita Orji"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18068","pdf_url":"https://arxiv.org/pdf/2609.18068","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM偏见","社会语言学","人类仿真"],"reason":"用LLM复现人类语言态度，有真实人类研究对照，揭示仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":12,"question":"大语言模型是否在住房相关的社会判断中继承了隐蔽的方言偏见，这种偏见是由种族身份、语言声望还是两者共同驱动的？","design":"采用匹配伪装范式，使用260组语义匹配的句子四元组（SAE、AAVE、NSE、NP），对十个开源权重LLM进行对数概率评分，测量模型在住房相关形容词上的概率差异，并设置三个社会距离不同的评估情境（租户筛选、邻居接受、室友选择）。","baseline":"人类基准来自社会语言学文献中记录的住房歧视研究（如Massey和Lundy 2001年的电话筛选实验）以及匹配伪装实验（如Tucker和Lambert 1969），但本研究未直接收集新的人类数据。","findings":"所有十个模型对AAVE和NP的输入一致分配更负面的住房相关形容词概率，其中NP受罚最重；每种方言通过不同的刻板印象簇被惩罚，而非作为泛化的非标准类别。NSE在正式租户筛选情境中比SAE更受青睐，但随着社会距离增加而受到惩罚，显示出情境依赖性。","reliability":"论文未明确讨论失效条件，但指出研究仅限于开源权重模型和特定方言，且未直接与人类判断进行定量对比，仅引用历史人类研究作为背景。","relevance":"该研究直接展示了LLM在住房筛选等经济相关场景中复现人类歧视性判断，并揭示了隐蔽概率偏见，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其匹配伪装设计，通过控制语义内容仅改变方言特征来隔离语言变量的因果效应，并利用对数概率测量隐蔽态度，可迁移到信贷审批中的方言歧视研究。｜可应用于金融领域的信贷审批或保险定价中的语言偏见问题，例如评估贷款申请人的方言口音是否影响模型的风险评估。｜设计：使用LLM作为虚拟信贷员，输入语义相同但方言不同的贷款申请文本（如SAE、AAVE、西班牙口音英语），测量模型分配给违约相关形容词或风险评分的概率，并与真实信贷审批数据中的种族或语言歧视率进行对照。"}},{"id":"2609.18274","version":1,"title":"I code or AI code: A comparative evaluation of AI-rated scores in classroom observations","zh_title":"我编码还是AI编码：课堂观察中AI评分与人类评分的比较评估","abstract":"Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GPT-5 model) to score teacher-child interactions in early childhood classrooms, benchmarked against human raters. The study analyzed 87 video-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System (CLASS) framework. AI-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning. Greater divergence emerged for interactions that were more procedural or context-dependent, particularly within the Classroom Organization and Instructional Support domains. These findings suggest that transcript-based AI scoring may capture some of the relative variation in teacher-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework. AI-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high-stakes evaluation. Future research should examine whether domain-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores.","authors":["Y. Fong","J. Xiang","T. Y. D. Chan","K. Lee","E. Y. H. Lau"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18274","pdf_url":"https://arxiv.org/pdf/2609.18274","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","教育评估","人类对照"],"reason":"用LLM替代人类评分员评估课堂互动，并与人类评分对照，属于仿真人类判断，但非典…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":13,"question":"使用大语言模型（GPT-5）对幼儿园课堂师生互动进行CLASS评分，能否替代人类评分员？","design":"用GPT-5模型基于课堂观察转录文本，按照CLASS框架对87个视频观察（来自香港30所幼儿园38个教室）进行评分，并与人类评分员在CLASS各领域和维度上的评分进行相关性和均值差异比较。","baseline":"人类评分员对相同视频观察的CLASS评分。","findings":"AI与人类评分在情感支持领域及反馈质量维度上收敛较好，但在课堂组织和教学支持领域等程序性、情境依赖性互动上分歧较大。转录文本AI评分能捕捉部分相对变异，但不能稳定复现人类校准判断。","reliability":"论文承认转录文本AI评分无法完全复现人类校准判断，建议仅作为初步筛查工具而非高利害评估替代品；未来需探索领域特定训练和纳入情境与视觉信息以改善对齐。","relevance":"该研究用LLM替代人类评分员评估课堂互动，并与人类评分对照，属于仿真人类判断，但非典型经济学实验或调查场景，与研究者关注的政策评估和经济学实验关联较弱，可作方法参考。","inspiration":"借鉴其用LLM对非结构化文本进行标准化量表评分的做法，并设置人类评分员作为基准进行相关性和均值差异检验。｜可迁移到经济金融中需要专家主观评分的场景，如信贷审批中的软信息评估、分析师报告语调分类、政策文本情感倾向等。｜以信贷审批为例，用LLM对贷款申请者的文字描述进行信用软信息评分，处理为不同提示词或模型版本，结果变量为信用评分，与人类信贷员的真实评分做对照，检验一致性和偏差。"}},{"id":"2609.18341","version":1,"title":"Understanding AI Provider Recommendations in Local Service Markets","zh_title":"理解本地服务市场中的AI提供商推荐","abstract":"When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.CY","cs.CL","cs.IR"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18341","pdf_url":"https://arxiv.org/pdf/2609.18341","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可靠性","审计研究","真实数据对照"],"reason":"审计LLM推荐与真实注册数据对照，揭示无检索时虚构与偏差，可迁移至仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":14,"question":"AI助手在本地服务市场中的提供商推荐是否真实、质量如何，以及检索配置如何影响推荐的可信度。","design":"审计研究：使用一个开源模型和一个专有模型（后者在有无网络搜索两种条件下）回答五类日常问询，覆盖美国100个最大都市区的四个注册服务领域，将推荐名称与官方注册数据（Medicare临床医生和设施记录、SEC顾问披露）进行匹配，测量匹配率、质量指标和偏差。","baseline":"官方注册数据：Medicare临床医生和设施记录（含星级评分）、SEC投资顾问披露（含不当行为记录）、餐厅众包评分和评论数。","findings":"无搜索时，模型在医生和养老院领域大量虚构推荐（匹配率仅4%-11%），且开源模型的匹配纯属姓名巧合；有搜索时匹配率升至64%-71%。搜索改变了推荐对象：无搜索时推荐的咨询公司SEC不当行为披露率是基准的3.6倍，有搜索时显著低于基准；搜索还消除了都市规模惩罚，使不同规模都市的匹配率趋于一致。","reliability":"论文指出，无检索的答案往往没有迹象表明推荐未经核实，且引用分析显示搜索推荐依赖商业来源而非监管来源（医生领域仅0.3%引用政府来源），存在可观察性差距；但未系统讨论模型幻觉的边界条件或不同提示措辞的影响。","relevance":"该研究通过对照真实注册数据审计LLM推荐，揭示了无检索时的虚构与偏差，可迁移至仿真可靠性评估，值得阅读原文以了解审计方法和偏差量化。","inspiration":"借鉴其审计设计：在同一模型内对比有无检索，用官方注册数据作为基准，测量匹配率和质量偏差，并分析偏差的稳健性。｜可迁移到金融顾问推荐、信贷产品推荐或医疗资源分配等场景，评估LLM仿真中的选择偏差。｜以LLM作为被试，处理为有无检索，结果变量为推荐实体的真实性和质量指标（如违规记录、评分），对照SEC或消费者金融保护局等官方数据，检验仿真是否复现或放大偏差。"}},{"id":"2609.18346","version":1,"title":"Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition","zh_title":"忠实却共谋：为何思维链监控无法检测寡头竞争下LLM定价代理的共谋","abstract":"Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination. We develop a causal graph divergence framework that separately measures structural faithfulness and intent faithfulness of LLM pricing agents in Bertrand competition. Across nine LLMs under duopoly and triopoly conditions, collusive behavior and chain-of-thought (CoT) faithfulness dissociate along both dimensions: the most collusive model accurately reports cooperative intent yet reasons structurally unfaithfully, while the most structurally faithful model sustains supra-Nash pricing under both market structures. These findings establish that CoT monitoring alone cannot serve as a standalone safeguard against algorithmic collusion.","authors":["Dohun Lee","Hyunwoo Park"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18346","pdf_url":"https://arxiv.org/pdf/2609.18346","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM定价代理","算法共谋","经济仿真"],"reason":"用LLM定价agent模拟寡头竞争，与人类行为对照，但非直接仿真人类被试，结论…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":15,"question":"LLM定价智能体在寡头竞争中能否通过思维链监控可靠地检测合谋行为？","design":"用9个LLM作为定价智能体，在双寡头和三寡头的伯特兰竞争环境中进行300轮定价，施加利润导向和竞争导向两种提示处理，测量定价序列、利润、思维链中的因果图与意图。","baseline":"无对照","findings":"合谋行为与思维链忠实性在两个维度上分离：最合谋的模型准确报告合作意图但结构推理不忠实，而结构最忠实的模型在两种市场结构下均维持超纳什定价。思维链监控不能单独作为算法合谋的保障。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体行为，虽非直接仿真人类被试，但涉及寡头定价与合谋检测，对关注LLM仿真可靠性及政策评估的研究者有参考价值，值得读原文了解其因果图分歧框架。","inspiration":"借鉴其因果图分歧框架，分别测量陈述与行为因果结构及意图分布，以评估LLM仿真忠实性｜可迁移到寡头定价、拍卖合谋、平台算法共谋等产业组织与反垄断场景｜用LLM作为定价智能体，施加不同提示（如利润导向vs竞争导向），测量定价序列与思维链，以真实市场定价数据或人类实验数据为基准，检验LLM合谋行为与思维链监控的有效性。"}},{"id":"2609.18390","version":1,"title":"Building a Cultural Perspective on Doctor-Patient Conversations","zh_title":"构建医患对话的文化视角","abstract":"AI-powered medical scribes are increasingly used to transcribe doctor-patient conversations and automate clinical documentation. However, large-scale real-world consultation datasets are scarce due to the sensitivity of clinical conversations, leading developers to rely on simulated and LLM-generated synthetic consultations. While scalable, these alternatives may fail to capture culturally situated patterns of clinical interaction. We introduce interactional cultural markers, measurable patterns of doctor-patient interaction grounded in cross-cultural clinical communication, and use them to compare real, simulated, and synthetic consultations from Indian and US clinical contexts. We find distinct patterns of participation and control: Indian consultations involve greater patient participation but stronger doctor control, while US consultations exhibit balanced participation and open-ended discussion. Synthetic Indian consultations often fail to reproduce these patterns, instead converging toward US-like interaction. We identify additional synthetic signatures, including excessive doctor explanation and formulaic patient responses. We conclude by discussing implications for generating culturally grounded synthetic clinical conversations.","authors":["Krithi Shailya","Siddharth D Jaiswal","Ashish Makani","Suvrankar Datta","Sunayana Sitaram","Mohit Jain"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18390","pdf_url":"https://arxiv.org/pdf/2609.18390","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","医患对话","文化差异"],"reason":"用LLM生成合成医患对话并与真实数据对照，评估文化模式复现，属仿真人类交互且含…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":16,"question":"不同来源（真实、模拟、合成）的医患对话在多大程度上再现了印度与美国临床互动中的文化模式？","design":"使用LLM生成合成医患对话（包括单智能体和多智能体架构，以及是否基于临床笔记的四种设置），并与真实和模拟对话进行比较，通过多层级互动文化标记（L1转录级、L2话轮间、L3话轮级）测量参与度、控制、语气等互动特征。","baseline":"来自印度和美国的真实医患咨询数据（包括英语、印地语、泰卢固语），以及由医疗专业人员或患者演员编写的模拟咨询。","findings":"印度真实咨询中患者参与度更高但医生控制更强，美国则更平衡且开放；合成印度对话往往无法再现这些模式，反而趋向美国式互动，并出现医生过度解释、患者公式化回应等合成特征。","reliability":"论文指出合成对话可能无法捕捉文化细微差别，且生成策略引入权衡：基于笔记的生成抑制口语和语码转换，而智能体生成则放大冗长和共情至不现实水平。","relevance":"该研究直接使用LLM生成合成对话并与真实人类数据对照，评估文化模式复现的可靠性，属于人类仿真研究，且包含批判性发现，值得阅读原文以了解其标记框架和失效条件。","inspiration":"值得借鉴的是其构建多层级可量化互动标记并系统比较真实与合成数据的方法，可迁移到经济金融中的跨文化沟通或谈判实验，例如不同文化背景下的信贷协商或政策沟通；具体设计可用LLM生成不同文化背景的谈判对话，以真实谈判记录为基准，测量话轮控制、信息共享等标记，检验合成数据是否复现文化差异。"}},{"id":"2609.18357","version":1,"title":"Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents","zh_title":"市场信号注入：对LLM定价智能体的对抗性情境操纵","abstract":"Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged. We introduce market signal injection (MSI), an attack that manipulates numerical formatting, competitor ordering, or qualitative market commentary without issuing explicit instructions. We evaluate nine open-weight models in simulated Bertrand duopoly and triopoly markets and three proprietary models in duopoly markets. Sentiment-based attacks produce the largest behavioral shifts, which propagate to other firms and alter profits and consumer surplus. Susceptibility varies across model families, and larger models are not consistently more robust. Matched neutral-text controls and a rule-based agent support a framing-based account of these shifts under the fixed demand parameters of our simulation. Episode-held-out probes distinguish baseline from attacked activations in all eleven re-evaluated model--condition pairs: linear AUC is 1.00 and MLP AUC ranges from 0.93 to 0.99. This separability does not by itself identify harmful pricing decisions. Input canonicalization removes the tested sentiment attacks, while decision boundary anchoring, which combines prompt constraints with output projection, provides partial mitigation under the tested adaptive attacks. These results identify data presentation as an attack surface for LLM pricing agents and motivate defenses that account for interactions among agents.","authors":["Dohun Lee","Hyunwoo Park"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18357","pdf_url":"https://arxiv.org/pdf/2609.18357","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM定价智能体","对抗攻击","市场模拟"],"reason":"LLM定价agent模拟市场，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.18394","version":1,"title":"Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland","zh_title":"语境中的文化能力：大语言模型在芬兰通过图灵测试","abstract":"We report the results of a Turing Test conducted in Finland in the Finnish language. Because languages and cultural contexts are unevenly represented in LLM training data, we expected the model (ChatGPT 5.2) to perform worse in a Finnish-language Turing Test than in previously studied English-language US contexts. We also present model-generated role prompting as a replicable technique for conducting comparative LLM-based Turing Tests designed to improve construct validity. Contrary to our expectations, the LLM passed the Finnish Turing Test. A prominent source of error was participants' reliance on linguistic cues, particularly colloquial Finnish, as markers of human authorship. We reframe the Turing Test from a test of intelligence to a comparative method for examining whether an AI system can display credible membership in a particular social world. Because its outcome reflects model capabilities, prompted identity, insider competence among human participants, and their AI literacy, the method provides a useful probe of the human-machine boundary across domains.","authors":["Otto Segersven","Pentti Henttonen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18394","pdf_url":"https://arxiv.org/pdf/2609.18394","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["图灵测试","文化能力","LLM评估"],"reason":"图灵测试评估LLM在芬兰文化中的表现，属于测量模型能力而非仿真人类被试，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.18591","version":1,"title":"Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making","zh_title":"递归推理还是统计外推？多智能体相互依赖决策中的上下文学习","abstract":"In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects refined internal reasoning or mere extrapolation of statistical patterns. To disentangle these mechanisms, we study LLM agents in multi-agent incomplete-information games that require recursive belief reasoning. By constructing a public goods game and manipulating the statistical structure of historical feedback, we evaluate decision quality against a history-independent rational expectations equilibrium (REE) benchmark. Our experiments reveal that when historical statistical patterns are disrupted, the benefits of longer context largely vanish, degrading decision quality to the no-context baseline in a way sharply amplified by stronger strategic interdependence. These results suggest that, in such strategic environments, ICL behavior is more consistent with statistical extrapolation than with strategic reasoning. Our work extends the mechanistic study of ICL to strategic multi-agent settings, introduces REE as a diagnostic tool for distinguishing reasoning from extrapolation, and provides a reusable framework for probing the boundaries of LLM reasoning in recursive belief tasks.","authors":["Yu Liu","Wenwen Li","Yifan Dou","Guangnan Ye"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18591","pdf_url":"https://arxiv.org/pdf/2609.18591","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","公共品博弈","社会模拟"],"reason":"用LLM agent模拟公共品博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.18384","version":1,"title":"GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models","zh_title":"GYROval：大语言模型文化价值取向的稳健基准","abstract":"We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYROval - Gridded Yielding of Robust value Orientation), together with the results of administering it to twenty models. Items are binary contrastive scenarios in the sense introduced by CDEval: both options are legitimate courses of action, neither is correct, there is no answer key, and a model's score on an axis is the proportion of its responses falling on the counted pole. Eleven of the twenty models were additionally administered a paired Russian translation of the identical items and a second sampling temperature. The instrument is publicly released in both languages. Stability was assessed by treating the vignette as the unit of analysis, ranking the models within the levels of each perturbation factor, and summarising the agreement between levels by tie-corrected Kendall's \\emph{W} against an empirical permutation null.","authors":["Alexander Didenko","Anna Shabanova","Vladislav Zapylikhin","Alexander Antipov","Ruslana Raemgulova"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18384","pdf_url":"https://arxiv.org/pdf/2609.18384","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["文化价值观","LLM测量","基准测试"],"reason":"测量LLM的文化价值观，属于把模型当测量对象，非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2608.24314","version":2,"title":"Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight","zh_title":"面向语音代理评估的LLM裁判基准测试：可靠性、校准与人工监督","abstract":"Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by comparing human judgments with GPT-4.1 and GPT-5 on telecom and retail voice-agent conversations, across conversational quality and safety dimensions. The same interac- tions are scored under three evaluation configurations, p0, p1, and p2, to test whether automated judgments are sensitive to the evaluation setup and whether observed patterns generalize across configurations and judge models. Beyond aggregate agreement, we examine metric-level correlations, evaluator consistency, and systematic human-LLM disagreement to identify which conver- sational attributes can be judged reliably by automation and which remain sensitive to interpretation and context. Effective voice-agent evaluation is also shaped by pipeline-level factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages, motivating our focus on comparing how human and LLM judges score the same interactions end to end. Our results show that LLM- based evaluation can serve as an effective component of large- scale voice-agent assessment, but that its reliability is metric- and configuration-dependent rather than uniform. This pro- vides an empirical framework for identifying which metrics suit automated evaluation and supports hybrid pipelines in which LLM judges handle scalable assessment while human evaluators remain engaged for metrics that demand contextual interpretation and higher-confidence judgment.","authors":["Anupam Purwar","Shashank Singh","Kritika Srivastava"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-08-26","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2608.24314","pdf_url":"https://arxiv.org/pdf/2608.24314","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","语音代理","人工监督"],"reason":"LLM作为评估者替代人工，但评估对象是语音代理而非人类被试，属于标注替代而非仿…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2608.25977","version":2,"title":"When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs","zh_title":"当人格遇上量化：量化LLM的逐层MBTI分析","abstract":"Personality is increasingly important in large language models (LLMs), as it shapes users' trust, engagement, and emotional experiences. While the Myers--Briggs Type Indicator (MBTI) has emerged as a common framework for assessing LLMs' personality, existing studies focus primarily on full-precision models and evaluate only final outputs. They overlook the widespread deployment of quantized LLMs requiring low memory footprints, whose personality traits remain underexplored. In this work, we present a systematic MBTI analysis of open-source LLMs across multiple precisions, including mainstream 4-bit methods (GPTQ, AWQ) and extreme 2-bit settings (AQLM variants). Beyond output-level evaluation, we examine how personality emerges across layers through option-level entropy and confidence-gap dynamics, and introduce Uncertainty-Amplified Layer Decoding (UALD) to study decoding-induced personality drift at inference time. Our results reveal a key insight: LLMs' personality is not a static property, but an emergent, layer-dependent decision process sensitive to quantization, prompting, and decoding. Specifically, we find that (1) ENFJ remains dominant across model families and precisions; (2) 4-bit quantization largely preserves coarse personality structure, while 2-bit quantization disrupts fine-grained prompt consistency and cross-precision agreement; (3) personality decisions emerges in upper layers, following substantial ambiguity in early layers; and (4) inference decoding can shift personality, while personality-aligned conditioning improves robustness. These findings provide a new perspective on the behavioral reliability of quantized LLMs and highlight the importance of considering internal dynamics and inference strategies in personality-sensitive chatbot applications.","authors":["Yao Fu","Lijia Huang","Xiaomin Li","Runchao Li","Yu Yin","Kenneth A. Loparo"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-17","first_seen":"2026-08-27","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2608.25977","pdf_url":"https://arxiv.org/pdf/2608.25977","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格","模型量化","MBTI"],"reason":"测量LLM自身人格，非仿真人类被试，但涉及人格测量与量化影响，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2609.18203","version":1,"title":"Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors","zh_title":"Behavior2Value：从电商行为中测量消费者价值观的基准与增强方法","abstract":"Human values are deep motivational orientations that shape human behaviors. In e-commerce, they reveal the stable drivers behind users' purchase decisions. Compared with short-term interests, consumer values better explain how users evaluate products before purchase. However, consumer values are often implicit in complex and fragmented behavioral trajectories, leaving value measurement from e-commerce behaviors largely underexplored. To this end, we propose the Behavior-to-Value (B2V) task, which aims to identify consumer values from e-commerce behavioral trajectories. Centered on this task, we first construct the E-commerce Consumption Value Taxonomy (ECVT) and introduce B2V-Bench, the first B2V dataset and benchmark, based on anonymized Taobao behavioral logs. B2V-Bench consists of real-world purchase decision episodes, covering 25 types of purchase behaviors, along with corresponding consumer value orientations manifested in each episode. To improve consumer value measurement accuracy, we further present B2V-Verifier, a behavior-to-value measurement model based on Value Verification Tuning, which learns to assess whether behaviors provide sufficient evidence for each value inference. Experiments show that B2V-Verifier outperforms strong LLM baselines, improving multi-label classification by 34\\%. The dataset and code will be publicly released upon acceptance.","authors":["Peixuan Hou","Bin Chen","Li He","Jian Xu","Bo Zheng","Xiuli Ma","Guojie Song"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18203","pdf_url":"https://arxiv.org/pdf/2609.18203","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM价值观测量","电商行为分析","基准数据集"],"reason":"用LLM从行为轨迹推断消费者价值观，测量对象是模型而非人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.18204","version":1,"title":"Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers","zh_title":"超越准确性：程序痕迹如何改变LLM监督者的决策标准","abstract":"Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 19 compliance tasks (4,551 analyzed judgments), varying only trace detail and evidence labeling. With disconfirming evidence always visible, error detection remains near ceiling. Instead, elaborate traces shift the decision criterion toward rejection, increasing false alarms in susceptible overseers. Without option labels, human-validated reason coding shows about 60% of false alarms cite an inability to tie evidence to its option. Labels eliminate this stated reason, yet residual rejection of correct work persists in those overseers and rises with trace detail. Procedural traces thus act as governance artifacts that shape oversight decisions. AI auditors should be evaluated by their decision criterion and false-alarm behavior, alongside accuracy.","authors":["Zihan Chen","Di Zhu","Lei Zheng","Weiling Li"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18204","pdf_url":"https://arxiv.org/pdf/2609.18204","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM审计","决策偏差","信号检测论"],"reason":"LLM作为审计者替代人工监督，非仿真人类被试，但涉及LLM决策偏差，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.17710","version":1,"title":"\"We Are Tired of Explaining\": Communication Practice and AI Roleplay Training for Community Health Workers in Rural India","zh_title":"“我们厌倦了解释”：印度农村社区卫生工作者的沟通实践与AI角色扮演培训","abstract":"Community health workers (CHWs) in the Global South increasingly encounter AI-powered tools, yet the counseling work central to their role remains largely unsupported. We study communication practices among Accredited Social Health Activists (ASHAs) in rural Rajasthan, India, through simulated family-planning calls, semi-structured interviews, and an LLM chatbot roleplay design-probe with 20 participants. In calls, ASHAs often responded to social or material concerns by shifting to health-risk information, denying concerns, promising unspecified help, or listing medical solutions with limited explanation. A smaller set of responses instead engaged concerns, sought permission before involving family members, or left decisions with beneficiaries. We interpret these patterns through Motivational Interviewing, emphasizing restraint from correcting, persuading, or over-solving. Drawing across observed calls, interviews, and probe reactions, we derive design considerations for AI roleplay training: keep AI in a rehearsal role, provide descriptive rather than prescriptive feedback, and evaluate counseling process rather than agreement with prescribed responses.","authors":["Neil K. R. Sehgal","Sunny Rai","Sai Preethi Matam","Khushboo Gupta","Hamid Abdullah","Mohit Jain","Sharath Chandra Guntuku"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17710","pdf_url":"https://arxiv.org/pdf/2609.17710","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI角色扮演","社区卫生工作者","培训工具"],"reason":"LLM用于角色扮演训练，替代人工陪练，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.18306","version":1,"title":"Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric","zh_title":"多智能体网络中的偏见放大：有偏智能体如何塑造观点与修辞","abstract":"Large language models (LLMs) are increasingly deployed in applications involving interaction between agents, where their output plays a role in collective reasoning and decision-making processes. Despite significant research into the functioning of LLMs in such multi-agent systems, the processes of bias propagation in such systems are still a challenge. This work studies how biased opinions are propagated in the form of textual interaction in an environment of LLMs, in which a minority of agents maintain persistent extreme opinions, while the remaining agents iteratively update their beliefs through structured textual interactions. The findings show that even the presence of a small percentage of biased agents in such a system leads to significant shifts in the opinions of non-biased agents. It suggests that for the same percentage of biased agents, the shifts occur more quickly for the Llama~3.2 model when compared to a classical Friedkin-Johnsen (FJ) model. Further semantic analysis demonstrates that rhetorical consistency in textual explanations increases systematically with biased exposure and, importantly, is partially decoupled from numerical convergenumericalutral agents adopt the vocabulary employed by the biased agents even in configurations where their numerical opinion shifts remain moderate. The research helps explain how bias and language develop together in multi-agent language model ecosystems.","authors":["Omran Berjawi","Giuseppe Fenza","Rida Khatoun"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18306","pdf_url":"https://arxiv.org/pdf/2609.18306","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体","舆论传播","偏见放大"],"reason":"多智能体舆论传播模拟，无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.18286","version":1,"title":"What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models","zh_title":"什么算战略推理？人类、引擎与语言模型国际象棋研究的系统映射","abstract":"Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence. The emergence of large language models (LLMs) has renewed the relevance of chess as a controlled environment for investigating strategic reasoning and comparing human and artificial decision-making. We present a systematic mapping study of recent research spanning human players, classical chess engines, neural and reinforcement-learning systems, LLMs, and hybrid approaches. The final map comprises 84 core study families, classified according to agent type, strategic-reasoning stages, and evaluation dimensions. The map reveals a literature strongly concentrated on situation assessment, evaluation, and action selection, while explicit planning, explanation, metacognition, and human--AI collaboration remain less explored. LLM research places particular emphasis on state representation and generalization, whereas grounded explanation appears more frequently in hybrid approaches combining language models with engines, expert knowledge, or other external structures. Two distinctions emerge that the map aggregates rather than resolves: hybrid systems differ in where and when heterogeneous capabilities combine, and evaluations that show improved human performance do not thereby establish human--AI synergy. We propose both as extensions of the mapping framework. We argue that chess provides a useful bridge between cognitive and computational perspectives on strategic reasoning, and identify explicit planning, grounded and faithful explanation, metacognitive calibration, and human--AI complementarity as directions for future research.","authors":["Paolo Ciancarini","Remo Pareschi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18286","pdf_url":"https://arxiv.org/pdf/2609.18286","source_feed":"cs.AI","score":4,"bucket":"other","rubric_hits":["C1"],"tags":["国际象棋","战略推理","系统映射"],"reason":"系统映射人类、引擎与LLM的国际象棋研究，非LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2609.12086","version":2,"title":"Creating an Atomic User Model for Personality-Aware Large Language Model Interaction","zh_title":"为感知人格的大语言模型交互创建原子用户模型","abstract":"Assistants built on large language models are expected to write in their users' own voice. Most systems summarise the user's preferences and include the summary in the prompt. This is the wrong way round. Preferences are only the surface of a person and change with the task, while the underlying personality stays the same, so storing preferences alone means relearning the user afresh whenever the task changes. This paper makes four contributions. First, we describe an effect we call personality seepage: the wording of a prompt carries traces of the writer's personality, which the assistant copies without knowing the writer. Second, we propose the Atomic User Model (AUM), a readable profile with a stable identity core surrounded by four layers covering psychological, cognitive, experiential, behavioral, and social details, plus notes on inner conflict and authenticity. Third, instead of inserting the entire profile, we use AUM as a searchable index, in which a task classifier, a selection step, and a budgeted retriever pass along only a few relevant fields. Fourth, we test the pipeline with 16 simulated users, 6 style-sensitive tasks, and 3 seeds. Eight retrieved fields matched the writing quality of the whole profile, while using only 23 percent of the context (211 tokens instead of 915). They scored 0.24 points higher than a plain preference note on a five-point scale. Accuracy in picking a user's own writing from four samples rose from 14.9 to 42.7 percent, where guessing gives 25 percent. Four pre-registered controls showed no effect, so the gain comes from the profile's structure rather than the search method. Personalization helps most for the users for whom a generic assistant imitates them the worst.","authors":["B. Sankar","Deepthika S","Pawni Yadav","Amogh A S"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-17","first_seen":"2026-09-14","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.12086","pdf_url":"https://arxiv.org/pdf/2609.12086","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["个性化助手","用户建模","写作风格模仿"],"reason":"论文聚焦个性化助手写作风格模仿，属于人格化聊天机器人，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.17536","version":1,"title":"Think Before You Comfort: Reflective Cognitive Alignment for Protocol-Grounded Elderly Stimulation Agents","zh_title":"安慰前先思考：面向协议约束的老年人刺激智能体的反思式认知对齐","abstract":"Cognitive Stimulation Therapy (CST) offers non-pharmacological support for elders with cognitive impairment, yet scalability remains constrained by reliance on trained facilitators and severe data scarcity, particularly for privacy-sensitive, low-resource languages such as Cantonese. While Large Language Models (LLMs) show promise for automated companionship, they often struggle to balance empathetic engagement with adherence to cognitive stimulation guidelines. We propose a framework addressing these challenges along two complementary axes. First, STaR-CS (Style-Transfer and Role-Conditioned Cognitive Stimulation) synthesizes multi-party dialogues through facilitator style modeling and structured skeleton extraction, mitigating data barriers. Building upon this corpus, the Reflective Cognitive Alignment (RCA) framework models stimulation interactions as a sequential decision process, integrating Protocol-Constrained Chain-of-Cognition (PC-CoC) for structured reasoning and Inference-Time Value Alignment (IVA) for principled response selection based on safety and engagement goals. Evaluations across six backbone LLMs and two independent judges show that RCA consistently improves protocol adherence, safety, and group facilitation over standard prompting baselines. Our code is available at https://github.com/jiangjyjy/RCA_Agent.","authors":["Jiyue Jiang","Ziyi Li","He Hu","Sheng Wang","Yuhan Chen","Yanyu Chen","Jingqi Zhou","Pengan Chen","Fei Ma","Irwin King","Yu Li","Chuan Wu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17536","pdf_url":"https://arxiv.org/pdf/2609.17536","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["LLM陪伴","认知刺激疗法","角色扮演"],"reason":"该研究用LLM做老年人认知刺激陪伴，属于角色扮演对话，无人类行为对照或仿真实验…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.18729","version":1,"title":"\"If I Had to Buy Just ONE: Galaxy S26 Ultra\": Auditing AI-Generated Product Recommendations","zh_title":"“如果我只能买一个：Galaxy S26 Ultra”：审计AI生成的产品推荐","abstract":"Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this raises difficult questions about the bias and impartiality of such advice. In response, we conduct an AI audit of popular chatbots using real commercial-advice queries. First, we curate a dataset of 2,528 real commercial-advice queries (ConsumerQ). Then, we evaluate 1,536 responses to product queries from popular AI chatbots: ChatGPT (chatbot and API), Google Gemini (chatbot and API), and Google Search (AI Overviews). We find that ChatGPT expresses a first-person product preference in 79% of product-recommending responses, compared with 7% for Gemini and 2% for AI Overviews, while the products recommended often change across repeated requests. Displayed sources vary strongly: for the same query, the ChatGPT and Gemini interfaces share only 5.4% of domains on average, with no domain in common in 76.7% of comparisons. APIs provide a different view from their corresponding interfaces, with mean domain overlaps of 12.0% for ChatGPT and 14.8% for Gemini, and also differ in the types and layers of source information they expose. Our findings show that neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter. Independent audits of AI-mediated commercial advice should therefore account for repeated responses, consumer-facing conditions, and the source layer being observed.","authors":["Lucas G. Uberti-Bona Marin","Thales Bertaglia","Giovanni Astante","Bram Rijsbosch","Gijs van Dijck","Anik\\'o Hann\\'ak","Gerasimos Spanakis","Konrad Kollnig"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18729","pdf_url":"https://arxiv.org/pdf/2609.18729","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI审计","产品推荐","偏见"],"reason":"审计AI产品推荐，属角色扮演对话，无人类行为对照或仿真目的","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.18998","version":1,"title":"One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs","zh_title":"单轴无刹车：自我知识限制LLM中有害同伴从众的过滤","abstract":"Multi-agent LLM systems are expected to be more reliable because agents can catch each other's mistakes. But peer pressure cuts both ways: the same correction that fixes a wrong answer can overturn a right one. The tempting safeguard is a brake that keeps the beneficial revisions and blocks the harmful ones. We show this brake is hard to build, for a simple reason: a revision is harmful exactly when the original answer was right, so deciding whether to block it is the same as knowing whether the model was already correct. This turns the open-ended hunt for a brake into one measurable quantity, the model's self-knowledge: any brake built from a deploy-time signal is a correctness probe in disguise, and self-knowledge is far from perfect (AUROC $\\approx 0.64$--$0.89$ across six model families). We call this ceiling the wall. Even white-box steering of the model's own correctness direction does not breach it: it changes how often the model revises, but harmful and beneficial revisions move together. At population scale the wall becomes the cliff: when most agents start wrong, debate amplifies the shared mistake into a confident, wrong consensus. In our multiple-choice societies, more agents, more model diversity, and a stronger member do not fix it. What helps is adding information before the revision, not filtering after it. Local agreement is not global correctness.","authors":["Yibo Hu"],"categories":["cs.LG","cs.CL","cs.MA"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18998","pdf_url":"https://arxiv.org/pdf/2609.18998","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","自我知识","共识形成"],"reason":"研究多智能体LLM的自我修正与共识，无人类行为对照，属纯多智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.11344","version":3,"title":"Governing Agentic AI in FinTech","zh_title":"金融科技中代理型人工智能的治理","abstract":"Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.","authors":["Henry Han"],"categories":["cs.CY","cs.AI","q-fin.RM"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-17","first_seen":"2026-08-13","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2608.11344","pdf_url":"https://arxiv.org/pdf/2608.11344","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","可验证性","金融科技"],"reason":"研究多智能体系统治理与可验证性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.14796","version":2,"title":"AI Persuasion as a Threat to Human Control","zh_title":"AI说服作为对人类控制的威胁","abstract":"The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.","authors":["Joshua Levy","Mick Yang","Kellin Pelrine"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-15","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.14796","pdf_url":"https://arxiv.org/pdf/2609.14796","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI安全","说服风险","风险评估"],"reason":"研究AI说服人类的风险，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.15293","version":2,"title":"Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures","zh_title":"为何LLM智能体在缺乏监督时会崩溃：执行差距作为涌现世界失败的机制","abstract":"When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced unanimous conformity -- without any external attacker. This paper identifies the mechanism. Reflexion-style agents already detect dangerous plan steps through iterative self-critique, yet the architecture provides no pathway from detection to action. We call this the enforcement gap: the audit sees the problem; the controller ignores it. Closing the gap requires a single conditional check -- fewer than 20 lines of code -- and reduces attack success by more than fourfold in large-scale experiments across frontier models, all five major agent frameworks, and an independent benchmark. We prove formally that when enforcement probability is near zero, detection quality is irrelevant to security. We further identify two compounding failure modes -- unreliable auditors and unparseable verdicts -- that explain every collapse pattern in Emergence World. A GRPO-trained enforcement controller resolves the ambiguity case. Concurrent work on filtering and information-flow control addresses the detection step but leaves the enforcement gap unaddressed; our results show this is the binding constraint. Together these results motivate a three-requirement Audit Enforcement Specification that is absent from every deployed framework today.","authors":["Yuhang Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-15","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.15293","pdf_url":"https://arxiv.org/pdf/2609.15293","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","安全机制","LLM智能体"],"reason":"多智能体仿真中的安全机制研究，无人类行为对照，不涉及人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.17301","version":2,"title":"When AI Becomes Hard to Understand: Cognitive Demands in Real-World Human-AI Conversations","zh_title":"当AI变得难以理解：真实世界人机对话中的认知需求","abstract":"Generative AI increasingly supports complex financial and health decisions, yet we know little about when its responses become difficult to process in real-world dialogue. We analyse more than 84,000 ChatGPT and Gemini conversations, using repeated prompting and clarification following misunderstanding as behavioural indicators of cognitive difficulty. We find that response characteristics such as length, readability and lexical diversity do not have fixed relationships with conversational difficulty; instead, their relationships depend on how they combine. Most notably, greater lexical diversity was associated with less repeated prompting in shorter responses, but this association weakened as response length increased, a pattern that replicated across financial and health conversations. We propose a conversational complexity budget to conceptualise these interdependencies: the demands associated with one response characteristic may depend on those accompanying it. The resulting design challenge is how to configure response complexity for the particular user, task and interaction.","authors":["Yingcan Carol Wang","Iman Munire Bilal","Qamar Zaman"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-16","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.17301","pdf_url":"https://arxiv.org/pdf/2609.17301","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","认知负荷","对话分析"],"reason":"研究真实人机对话中的认知难度，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2609.17552","version":1,"title":"Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks","zh_title":"道德推理训练有益还是有害？用角色扮演攻击对RL训练的伦理智能体进行红队测试","abstract":"Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such attacks are realistic: retrieved context, tool outputs, or multi-turn framing can all inject role instructions that compete with the agent's moral objective. We red-team morally trained Gemma-2-27B/9B and Llama-3.1-8B agents with five persona attacks, then probe causality with noise-reward controls, adversarial PPO, representation analysis, steering, and head ablations. At 27B, moral RL cuts mean adversarial degradation by 5.2x but costs ~11pp ETHICS accuracy; across 205 scenarios and 5 seeds, reasoning-level moral reward yields 5.8x robustness while a matched random reward yields none. The training also reshapes representation geometry (mean CKA 0.82/0.83 vs. 0.98 for noise), moves peak attack processing 8 layers earlier, and exposes a rank-1 L21 direction that recovers 83% of full PPO's average robustness. One failure mode survives all of this. Against Fiction role-play, L21 steering recovers only 29% of the gap, and head ablation finds 38 compliance heads competing with 25 alignment heads. Moral RL thus builds robustness that is partly linear and partly circuit-distributed, transferable through activation steering, yet still beaten by named-character role-play.","authors":["Arth Singh"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17552","pdf_url":"https://arxiv.org/pdf/2609.17552","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM安全","角色扮演攻击","道德对齐"],"reason":"研究对抗性角色扮演对道德训练LLM的影响，属角色扮演攻击鲁棒性，无人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.17853","version":1,"title":"AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content","zh_title":"AfriSyCo：测量非洲语言内容中的断言框架、验证与措辞敏感性","abstract":"AfriSyCo studies answer switching around African-language factual content with two complementary layers: native-language follow-ups and a controlled cross-language factorial whose question, options, and target remain in the African language while the follow-up framing is English. We analyze 1,415 turn-1-correct model-language-item observations derived from 100 source questions across seven open-weight checkpoints and six languages; turn-1-correct denotes observed first-response accuracy, not demonstrated knowledge. Under native prompts, assertive endorsement produces 29.3 percentage points more any-turn false-target selection than mention-plus-verification (M+V), with a 19.0-point immediate T2 contrast. In the precommitted 2 x 2 factorial, averaged over three tested prompt families, assertive framing increases target selection by 30.4 points (95% CI [28.4, 32.3]); verification decreases it by 17.4 points, while the assertive effect rises from 20.5 points without verification to 40.2 with it (interaction +19.7). The effect remains 34.8 points among 611 observations correct after option reordering. Magnitude varies sharply by wording and checkpoint: prompt-family effects span 20.1-42.5 points, a Twi/Qwen3 paraphrase shifts target selection from 70.8% to 4.2%, and checkpoint effects span 9.2-47.0 points. Prompt realization is therefore part of the measurement problem.","authors":["David Ababio Awuni","Rose-Mary Owusuaa Mensah Gyening","Elvis Gyasi Owusu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17853","pdf_url":"https://arxiv.org/pdf/2609.17853","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","非洲语言","回答切换"],"reason":"研究LLM对非洲语言事实内容的回答切换，属纯NLP能力评测，无人类被试仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.17857","version":1,"title":"Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels","zh_title":"谁当裁判很重要：测量LLM裁判面板中的家族条件偏好","abstract":"Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confounded with candidate quality and correlates with Bradley-Terry ability at r = 0.95. We derive a corrected estimator that holds the candidate family fixed and compares judges. All four families then show a positive same-family lift (3.4-8.4 percentage points), with global FPS 0.067 (95% CI [0.053, 0.084], permutation p = 0.0002). The effect remains under panel-based quality controls, an independent human-consensus anchor, and a float16 judging replication. Judge-side likelihood is closely related to the effect: adding likelihood advantage reduces the controlled coefficient by 61%, which we treat as descriptive attenuation rather than causal mediation. Position is a separate failure mode. Across the panel, 55.4% of AB/BA pairs reverse, and reversal above 50% is incompatible with a simple independent content-noise model. Relative to a family-balanced reference, panel composition changes 18.5% of pairwise outcomes. A complete reproducibility archive has been prepared for public release.","authors":["David Ababio Awuni","Luke E. K. Achenie","Benjamin Tei Partey","Elvis Gyasi Owusu","Nii-Nai Derrick Sowah"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17857","pdf_url":"https://arxiv.org/pdf/2609.17857","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评估","裁判偏差","多智能体"],"reason":"研究LLM作为评判者的偏好，不涉及人类被试仿真或人类行为对照，属于多智能体评估…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.18649","version":1,"title":"DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions","zh_title":"DyMT-ESB：用户与LLM交互中社会偏见的动态多轮评估","abstract":"Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios important for safety, including stereotyping-related harms. However, existing multi-turn social bias evaluations often rely on pre-specified or template-based user inputs that do not adapt to model responses and typically assume a fixed dialogue length in advance. In this paper, we study social bias dynamics in response-conditioned multi-turn interactions using a controlled evaluation protocol that generates follow-up user queries from the evolving dialogue history and allows evaluation over variable numbers of turns. Experimental results show that LLMs exhibit social bias even in coherent, response-conditioned multi-turn interactions, revealing late-emerging bias, non-monotonic bias patterns, and bias re-emergence. These results motivate evaluations that extend beyond fixed-turn, pre-scripted protocols. Our findings highlight the importance of analyzing social bias as a turn-level dynamic phenomenon.","authors":["Rem Hida","Masahiro Kaneko","Daisuke Oba","Danushka Bollegala","Naoaki Okazaki"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18649","pdf_url":"https://arxiv.org/pdf/2609.18649","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["社会偏见","多轮对话","模型评测"],"reason":"评估LLM在多轮对话中的社会偏见，属于模型安全评测，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.18905","version":1,"title":"Structured Claim-Level Discourse Representations for Dense Health Narratives","zh_title":"密集健康叙事中的结构化声明级话语表示","abstract":"Health discourse in social media videos often contains densely entangled claims spanning multiple thematic aspects, stances, evidential frames, and rhetorical functions within short conversational spans. Existing approaches largely rely on coarse topic-level, sentiment-based, or stance-oriented representations that do not adequately capture this structure. Our analysis identifies an average of 13.22 atomic claims per minute, motivating richer claim-level discourse representations. We introduce a structured framework for claim-level discourse analysis in dense health narratives. Our framework models discourse through tuples linking atomic claims with thematic aspects, stance, and multidimensional pragmatic discourse attributes. To support this setting, we construct a benchmark spanning four health domains with 1,191 manually annotated claims from 60 videos. Using this framework, we evaluate automated structured discourse analysis under different discourse context settings. Results show that current LLMs achieve strong performance on thematic categorization and stance prediction, but struggle with high-dimensional pragmatic profiling. We also find that different discourse tasks benefit from different forms of contextual reasoning, suggesting that future systems may require task decomposition and specialized inference strategies.","authors":["Farnoushsadat Nilizadeh","Elham Pourabbas Vafa","Shirin Nilizadeh","Eduard Dragut"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18905","pdf_url":"https://arxiv.org/pdf/2609.18905","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["话语分析","健康叙事","LLM评测"],"reason":"论文是NLP结构化话语分析评测，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.18908","version":1,"title":"How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards","zh_title":"人权价值几何？ECtHR-NPD：预测非金钱损害赔偿金的基准","abstract":"Existing legal benchmarks cover diverse tasks, while continuous monetary remedies remain comparatively underexplored. We introduce ECtHR-NPD, to the best of our knowledge, the first benchmark for predicting non-pecuniary damage (NPD) awards at the European Court of Human Rights (ECtHR) from case information when no statutory formula or explicit calculation rule determines the amount. ECtHR-NPD contains 14,575 cases with case-level awards in nominal euros, chronological splits, and a protocol separating target construction from model input. We evaluate a battery of methods, including constant predictors, gradient-boosted trees, retrieval methods, fine-tuned encoder language models (LMs), prompted decoder LMs, and knowledge-augmented agents. Our results show that more sophisticated LM and agentic approaches do not consistently outperform the strongest feature-based baseline. All model families struggle to identify zero awards and to calibrate high-award predictions, with further degradation on the Challenging test view, making ECtHR-NPD a challenging testbed for current state-of-the-art open-weight and proprietary LMs.","authors":["Yanyi Pu","Damian A. Gonzalez-Salzberg","Zheng Yuan","Nikolaos Aletras"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18908","pdf_url":"https://arxiv.org/pdf/2609.18908","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["法律NLP","损害赔偿预测","基准测试"],"reason":"法律判决金额预测，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.18960","version":1,"title":"When Audit Quality Fails to Predict Downstream Utility: A Counterfactual Study of Synthetic-Data Selectors for Low-Resource African NLP","zh_title":"当审计质量无法预测下游效用：低资源非洲NLP合成数据选择器的反事实研究","abstract":"Quality-aware synthetic-data selection rests on a proxy: examples that an LLM judge rates as good should also help a downstream model learn. In a controlled replay in low-resource African-language classification, we show that this proxy breaks. Across four languages (Amharic, Hausa, Swahili, Yoruba), two classification tasks (MasakhaNEWS, AfriSenti), and five matched-budget selectors, audit rankings and downstream rankings diverge. Within each cell, the Spearman between judged label correctness and Macro-F1 across selectors has mean $\\rho{=}0.04$ (median $0.00$), showing that the mismatch is not an aggregation artifact. \\method{}-V2, our counterfactual audit framework, produces the cleanest selected pool on three audit channels at once: highest judged label correctness ($0.904$ vs.\\ $0.767$ for naive, a $17.9\\%$ relative gain), lowest shortcut score, and a hard-reject rate of $0.162$ vs.\\ $0.486$ for naive. AlpaGasus nevertheless leads downstream Macro-F1 ($0.202$ vs.\\ $0.163$ for \\method{}-V2), and the inversion persists on the five non-degenerate cells. The lesson is methodological: in this controlled setting, audit quality is a property of the selected pool, not a guarantee of downstream utility. Synthetic-data evaluation should therefore report audit and downstream metrics on the same retained sets. We release the audit tables, per-selector retained pools, and a claim ledger that links every reported number to its source row.","authors":["Son Ha Xuan","Phat T. Tran-Truong","Xuan-Bach Le"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18960","pdf_url":"https://arxiv.org/pdf/2609.18960","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["合成数据","数据选择","低资源NLP"],"reason":"论文研究合成数据选择器的审计质量与下游效用，不涉及人类行为仿真或对照，属于NL…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2609.19006","version":1,"title":"WordPolo: Evaluating Language Models Through Iterative Semantic Feedback","zh_title":"WordPolo：通过迭代语义反馈评估语言模型","abstract":"Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their reasoning processes. We present WordPolo, a word-finding task where participants must discover an unknown target word using semantic similarity feedback. Players start with zero knowledge, make guesses, and receive distance scores (1 = correct, higher = further away). Success requires interpreting scores to navigate semantic space and systematically narrow the search. This design makes iterative reasoning and adaptive search strategies both directly observable and necessary for success. We evaluate recent LLMs (GPT-4.1, Llama 4, Claude 3.5 Haiku, Qwen 3), LRMs (o4-mini, Deepseek-R1), humans, and a novel heuristic on 1,500 puzzles. Beyond solve rates (which range from 4% to 62%), we introduce progression-based metrics that reveal models often make meaningful progress, insights that accuracy alone would miss. Our analysis shows how reasoning models can be hindered by overthinking and underthinking, while successful models exhibit human-like strategies. WordPolo demonstrates the need for benchmarks that test both reasoning process and outcomes, providing holistic measurements of model capabilities. Our code and dataset can be found at https://wordpolo-demo.vercel.app/.","authors":["Tyler McDonald","Ali Emami"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19006","pdf_url":"https://arxiv.org/pdf/2609.19006","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","推理过程","语义搜索"],"reason":"纯LLM能力评测，虽含人类对照但非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2609.17632","version":1,"title":"EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents","zh_title":"EvolveTrade：自进化LLM交易智能体的经验驱动策略精炼","abstract":"Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk under changing market regimes. We introduce EvolveTrade, a self-evolving framework that treats the system prompt of a tool-using trading agent as a text-parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and realized portfolio feedback, while keeping the backbone LLM fixed. The updated policy is then used for the next batch of trading decisions, enabling the agent to refine its information-acquisition and portfolio-construction procedure over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines, achieving the improved SR and CR in most evaluated settings. Behavioral analyses further show that self-evolved policies increase code-mediated analysis and activate regime-relevant computations; case-level policy-to-return attributions trace how policy-induced allocation changes contribute to realized return differences. These results suggest that adapting the reusable procedure governing tool use is a key direction for building more robust LLM trading agents.","authors":["Sehee Kim","Yumin Choi","Minki Kang","Sung Ju Hwang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17632","pdf_url":"https://arxiv.org/pdf/2609.17632","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","金融交易","策略优化"],"reason":"LLM交易智能体自我进化，优化工具使用策略，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.17847","version":1,"title":"Learning Heterogeneous Preferences","zh_title":"学习异质性偏好","abstract":"Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \\emph{universal utility} function shared across a population and treat disagreement between annotators as stochastic variation. While suitable for objective tasks, this assumption breaks down in subjective domains where preferences vary systematically across individuals. We study the problem of subjective preference learning, in which observed choices arise from heterogeneous but internally consistent utility functions. Drawing upon rational choice theory, RCT \\parencite{tversky1981framing}, we introduce \\emph{individuated utility} functions conditioned on both the individual and their decision context, and propose a novel multi-stage architecture for estimating them from multi-modal data. We evaluate our framework on a newly collected dataset of more than $575{,}000$ pairwise aesthetic judgments from $2{,}398$ participants comparing automotive wheel designs. Our experiments show that individuated utility models substantially outperform universal utility models including foundation model baselines. Our results demonstrate that disagreement reflects meaningful preference heterogeneity rather than annotation noise. More broadly, our findings highlight the importance of collecting annotator attributes and learning individuated utility functions, enabling reward models that explicitly account for whose preferences they represent and faithfully capture human decision diversity.","authors":["Shiwali Mohan","Matt Hong","Dule Shu","Aniek Fransen","Shabnam Hakimi","Matt Klenk"],"categories":["cs.AI","cs.HC","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17847","pdf_url":"https://arxiv.org/pdf/2609.17847","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["偏好学习","人类反馈","个性化建模"],"reason":"论文用人类偏好数据训练模型，而非用LLM仿真人类被试，方向相反。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.17772","version":1,"title":"When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments","zh_title":"当AI生成协变量：序贯实验中的因果类型与估计目标漂移","abstract":"AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.","authors":["Takes Fujita (VRI)","Nobutaka Hattori (Department of Neurology, Juntendo University School of Medicine)"],"categories":["stat.ME","cs.AI"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17772","pdf_url":"https://arxiv.org/pdf/2609.17772","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["因果推断","AI生成协变量","统计方法"],"reason":"论文讨论AI生成协变量在因果推断中的角色，不涉及用LLM仿真人类被试或与人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.17883","version":1,"title":"Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming","zh_title":"AI辅助是否留下时间指纹？检测AI辅助写作与编程中的过度依赖","abstract":"The rapid adoption of generative AI has made final artifacts unreliable evidence of student learning, and AI detectors that examine only the finished product are inaccurate and ethically contentious. Process data offers an alternative, but prior work covers only English essay writing. We ask whether AI assistance carries a temporal signature, whether it generalizes from writing to programming, and whether it distinguishes ordinary collaboration from wholesale delegation. We analyze three public corpora: CoAuthor (1,447 keystroke-level co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus (5.1 million keystrokes) as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation. Three findings emerge. First, the signature generalizes: AI contributions arrive in bursts far outside the author's own baseline in both mediums (paired d_z = 1.13 and 3.54). Second, engagement diverges by medium: 93% of AI-inserted characters survived to writers' final documents, while only 14% of accepted code suggestions survived intact. Third, classifiers using only observable temporal features separate simulated delegation from authentic work nearly perfectly (F1 $\\geq$ 0.997; at most 0.5% of real work misclassified), while ordinary collaboration remains hard to distinguish from unassisted work. Temporal evidence flags wholesale delegation rather than assistance, positioning process visibility as a candidate evidentiary basis for academic integrity, pending validation in authentic coursework.","authors":["Eduardo Davalos","Yike Zhang"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17883","pdf_url":"https://arxiv.org/pdf/2609.17883","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助检测","学术诚信","过程数据"],"reason":"研究AI辅助写作与编程中的过度依赖检测，不涉及用LLM仿真人类被试或与人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.18949","version":1,"title":"StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction","zh_title":"StableEval Arena：面向稳定币价格稳定性预测的成本感知智能体基准","abstract":"We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed agentic systems on diagnosing peg stress and forecasting deviations from the one-dollar peg over a hidden seven-day horizon, using leakage-safe historical replay with exchange price-volume data and market-context features. We report two complementary experiment blocks: a 120-case stress-enriched validation block and a 507-case natural-distribution full-arena evaluation block. Across six LLM-backed agent configurations and baselines, StableEval Arena measures prediction quality, calibrated-label behavior, structured-output reliability, latency, token consumption, and estimated inference cost. Rather than ranking agents by accuracy alone, the framework treats trustworthiness as a joint property of forecast quality, operational reliability, and computational cost. The results show a gap between protocol-following reliability and financial-risk reliability: agents reliably produce valid structured outputs at modest measured cost, but still miss most rare severe-stress and sustained-depeg cases. To support auditing and replication, we release the benchmark dataset on Hugging Face and the source code on GitHub.","authors":["Sean Wan","Dongping Liu","Luyao Zhang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18949","pdf_url":"https://arxiv.org/pdf/2609.18949","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","金融预测","基准测试"],"reason":"评估LLM智能体预测稳定币价格，属金融预测基准，非人类行为仿真，无人类被试对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.17948","version":1,"title":"Apply-<x>Mag: One Tool to Support Many Inclusive Design Methods","zh_title":"Apply-<x>Mag：一个支持多种包容性设计方法的工具","abstract":"Doing inclusive design in HCI practice can be labor-intensive, a costly barrier that some companies and HCI practitioners may be unwilling or unable to overcome. Yet, not doing inclusive design is costly too, in the form of UX barriers that disproportionately disadvantage under-served user populations. To address this problem, we introduce Apply-<x>Mag, an LLM-powered tool to support HCI practitioners' work to design their products inclusively to wide ranges of users. Apply-<x>Mag is general, supporting any inclusive design method that can be expressed as <x>Mags (i.e., using attribute ranges and heuristics). It is also effective: Empirical results with researcher and practitioner teams using various combinations of two <x>Mags on 7 products showed Apply-<x>Mag precision averaging 90-99% and recall averaging 82-89%. Further, its environmental costs were reasonable, costing about the same resources as 2-4 ordinary Google searches.","authors":["Sadia Afroz","Rudrajit Choudhuri","Fatima A. Moussaoui","Amreeta Chatterjee","Margaret Burnett","Anita Sarma"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17948","pdf_url":"https://arxiv.org/pdf/2609.17948","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM工具","包容性设计","HCI"],"reason":"LLM用于辅助包容性设计，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.18479","version":1,"title":"Verify, Offload, Extend & Recommend: Selective Complementarity in AI Support for Physical Activity Planning with Longitudinal Patient Data","zh_title":"验证、卸载、扩展与推荐：纵向患者数据下体力活动规划中AI支持的选择性互补","abstract":"Self-tracking technologies create longitudinal patient-generated health data, yet integrating these data into clinical decision-making can increase information-processing demands. Generative AI may support sensemaking, but its value depends on clinical context and expertise. We investigate AI augmentation of a clinical decision support system for physical-activity planning in cardiovascular disease. In a counterbalanced within-subjects study, 26 exercise physiologists developed plans for four real cardiovascular cases with and without AI support, followed by evaluation of an AI exercise-plan generator. AI did not significantly improve workload, usability, confidence, or plan quality overall; however, its effect on plan quality increased as visualization literacy decreased and its effect on workload increased as visualisation literacy increased. Interviews and 152 chatbot queries revealed three recurring uses: verifying, offloading, extend and generate. Our findings position AI support as a selective complement to professional expertise while highlighting validation challenges when clinicians seek support precisely where their own knowledge is limited.","authors":["Pavithren V S Pakianathan","Rania Islambouli","Diogo Branco","Gil Batista Rosa","Rita Pinto","Albrecht Schmidt","Tiago Guerreiro","Jan David Smeddinck"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18479","pdf_url":"https://arxiv.org/pdf/2609.18479","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助决策","临床支持","人机交互"],"reason":"研究AI辅助临床决策，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.17310","version":2,"title":"Zero-shot narrative detection in social messaging","zh_title":"社交消息中的零样本叙事检测","abstract":"This study investigates the zero-shot ability of large language models (LLMs) to identify and classify hidden narratives in social messages. Our research hypothesis is that LLMs' extensive contextual knowledge allows them to interpret messages on a deeper, pragmatic level, going beyond basic sentiment or topic analysis. Experiments on the Dipromats and SemEval datasets show that providing models with human-written narrative descriptions significantly improves performance, without the need of training examples. In contrast, automatically generated descriptions or the use of few examples (few-shot) often degrade accuracy due to subtle shifts in framing. The study also finds that ensemble methods, particularly majority voting, enhance robustness and that larger models perform best while also being less sensitive to prompt variations. The findings validate that LLMs can effectively detect strategic narratives in a zero-shot setting, and when combined with simple ensembling and human-written descriptions, they can rival supervised systems, offering a scalable solution for narrative detection, specially when there is no training data for the vast majority of domains.","authors":["Jes\\'us M. Fraile-Hern\\'andez","Anselmo Pe\\~nas","Patrick Giedemann"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-16","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.17310","pdf_url":"https://arxiv.org/pdf/2609.17310","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["叙事检测","零样本学习","文本分类"],"reason":"纯NLP能力评测，检测叙事而非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2609.17532","version":1,"title":"Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes","zh_title":"利用呼吸治疗临床笔记的LLM衍生特征增强拔管失败预测","abstract":"Invasive mechanical ventilation is a lifesaving therapy, but timely, safe discontinuation is essential to preventing extubation failure (EF) and related risks to health. We present a novel approach to EF prediction that leverages features classified in free-text respiratory therapy notes using a large language model and logistic regression pipeline. Applied to a patient cohort from University of Washington Medicine, our method identifies clinically meaningful EF-related features that improve EF prediction performance when included alongside structured patient data. We further highlight how differences in target populations in prior EF prediction studies, such as heterogenous inclusion criteria and EF definition, can lead to systematic differences in model performance and hinder generalizability between studies.","authors":["Izzy Chaiken","Aditya Khowal","Neha A. Sathe","Mark M. Wurfel","Lucy Lu Wang"],"categories":["cs.CL","cs.LG","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17532","pdf_url":"https://arxiv.org/pdf/2609.17532","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床预测","LLM特征提取","医疗NLP"],"reason":"论文用LLM提取临床特征以改进拔管失败预测，属于医疗NLP应用，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.17602","version":1,"title":"Making Political Text Scaling Comparable: Infrastructure and Hyperparameter Sensitivity for 17 Algorithms","zh_title":"使政治文本尺度化具有可比性：17种算法的基础设施与超参数敏感性","abstract":"Computational text-based ideal point estimation (CT-IPE) methods are usually compared as named algorithms, yet applying them involves numerous researcher choices that configure how political text is turned into position estimates. This paper argues that CT-IPE methods are better understood as configurable measurement pipelines than as fixed estimators. Building on a large-scale comparative experiment spanning 17 CT-IPE algorithms, 5,537 experimental runs, and approximately 4.25 million left-right position estimates, I describe the shared infrastructure that makes these heterogeneous methods jointly executable and quantify how sensitive their estimates are to alternative hyperparameter choices. Variance-partitioning and SHAP-based sensitivity analyses show that, for most algorithms, hyperparameter profiles explain little residual variance through a shared shift: 13 of the 17 algorithms exhibit ICC values below .10. Where this profile-level sensitivity is present, it is concentrated in a small number of consequential researcher choices, most notably the selection of the underlying language or embedding model, the seed keyword lists that anchor the construct, and the number of topics.","authors":["Patrick Parschan"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17602","pdf_url":"https://arxiv.org/pdf/2609.17602","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文本尺度化","算法比较","政治文本"],"reason":"论文比较文本尺度化算法，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.18385","version":1,"title":"Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts","zh_title":"多模态社交媒体帖子中的情感体验、表达与感知","abstract":"Emotions are an essential aspect of human communication, particularly on social media, where authors frequently combine text and images to convey their emotions. Yet prior work on emotion analysis of social media posts has overlooked two important aspects in regard to measuring how well readers can reconstruct the authors' intent: (1)~the image modality, with most work focusing solely on text, and (2)~the real-world events that trigger the expressed emotions, and their relationship to the post content. We therefore study the relation between (a) the author's experience of the event that caused them to write a social media post and (b) the content of the post, with a focus on readers' capability to reconstruct that emotion expression. To do that, we introduce the Multimodal Multi-Emotion-Model dataset Mult2EMo, created by collecting annotations from both authors and readers on the posts and their triggering events. We find that reconstruction is possible but challenging for both human readers and computational models. We show that understanding the triggering event is crucial for accurate reconstruction, and that reconstruction is particularly challenging when posts rely heavily on the image to express emotion.","authors":["Christopher Bagdon","Carina Silberer","Roman Klinger"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18385","pdf_url":"https://arxiv.org/pdf/2609.18385","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","多模态","数据集"],"reason":"纯情感分析数据集与模型评测，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.18272","version":1,"title":"Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI","zh_title":"谁审计谁，在什么基座上，用什么证据？面向智能体AI的独立性分级审计协议","abstract":"Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor. Independence, the foundation of assurance,is still applied to them as a binary. We argue that it must be graded along three orthogonal axes: principal independence (who controls the auditor), substrate independence (an auditor sharing the auditee's foundation-model family, toolchain or guardrails fails with it) and evidence independence (whether evidence is attestable rather than self-reported). Each axis has precedent; the contribution is to grade all three on a single audit, aggregate them by the weakest link, and apply the same rubric when the auditor is itself an agent. We give the model a formal basis by transplanting the beta-factor model of common-cause failure from reliability engineering, a seven-step protocol whose outputs a third party can verify, a structural detectability analysis of a procurement-controls agent audited at three grades, and a Monte Carlo study of the model in which a conventional internal audit of an agent-a real audit team, a second agent, provider logsp-surfaces 5.9% of the faults it could in principle see and none at all in half the fault classes. We map the triple to the EU AI Act as amended, ISO/IEC 42006, UK public-sector risk-management guidance and audit-regulator practice.","authors":["Mohamed Chahine Ghanem"],"categories":["cs.AI","cs.CR","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18272","pdf_url":"https://arxiv.org/pdf/2609.18272","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI审计","多智能体系统","独立性评估"],"reason":"论文研究AI审计协议，涉及多智能体协作但无人类行为仿真或对照，属于纯多智能体系…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.18676","version":1,"title":"The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses","zh_title":"生成式AI对学生学习的不均衡影响：考察AI相关课程中依赖、评估素养和课程政策的作用","abstract":"Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance remain underexplored. Using survey responses from 118 students across 12 AI-related courses at our institution, we examined differences in GenAI use and perceived learning experiences. We identified four user clusters: high-use students reporting many benefits, light users reporting less reliance and fewer benefits, and two moderate-use groups reporting different levels of benefit. We also found significant differences between free- and premium-version users, single- and multiple-tool users, and students experiencing different instructor policies. In multivariable regression models, academic benefit was associated with early reliance and academic task support; positive impact was associated with cognitive reliance, academic task support, confidence in GenAI reliability, and instructor policy; and negative impact was associated with early reliance and attitudinal change. The association between early reliance and negative impact became stronger as evaluation literacy increased. Finally, perceptions of GenAI-enhanced learning appear to reflect cognitive, performance, and self-efficacy benefits, while concerns about stress and diminished critical thinking are associated with lower perceived learning benefits. These findings suggest that institutions need better policies to address such inequities so that institutions can enable students to benefit from increasingly capable AI systems.","authors":["Lydia Manikonda","Mei Si","Sirajam Munira","Oshani Seneviratne","Kristin Bennett"],"categories":["cs.AI","cs.ET","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18676","pdf_url":"https://arxiv.org/pdf/2609.18676","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["教育技术","学生调查","GenAI使用"],"reason":"研究人类学生使用GenAI的学习效果，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.17779","version":1,"title":"AI and Human Approaches to Mathematical Problem Solving","zh_title":"AI与人类解决数学问题的方法比较","abstract":"AI systems have begun to report solutions, disproofs, and substantive advances on long-standing mathematical problems, raising questions about whether they approach research in the same way as mathematicians. This study compares public AI research accounts with the human literature on 11 such problems. The human corpus contains 58 papers that directly addressed the same mathematical targets later reported by AI sources as resolved, disproved, or substantially advanced; 31 within-problem comparisons were constructed from these materials. Six validated text-based measures capture problem resolution, method articulation, uncertainty and boundary specification, successor-question generation, generality, and cross-disciplinary integration. AI accounts place greater emphasis on resolving the focal problem and connecting ideas across fields. Human papers devote significantly more attention to explaining methods, specifying assumptions and limitations, and identifying questions for subsequent research. No precise difference is detected in generality. The estimated directions remain unchanged when each mathematical problem is removed in turn. The findings reveal two distinct research profiles: AI accounts concentrate on closing and recombining problems, whereas mathematical papers more extensively document the procedures, limits, and research opportunities through which results become cumulative knowledge. Evaluating research AI therefore requires attention to the organization of inquiry, not only whether a target is solved.","authors":["Yang Ding"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17779","pdf_url":"https://arxiv.org/pdf/2609.17779","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI研究行为","数学问题解决","文本分析"],"reason":"比较AI与人类数学研究风格，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.18505","version":1,"title":"Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course","zh_title":"整合翻转学习与生成式AI的实践型设计教育：来自针织纱线设计课程的证据","abstract":"In practice-based design courses such as knit yarn design, students must turn visual ideas into feasible material outcomes. This is difficult because creative decisions are tied to yarn properties, stitch structures, machine operation, and limited opportunities for physical sampling. This study presents an integrated pedagogical framework that combines flipped learning, exemplar-based reference, GenAI-assisted visual prototyping, and studio feedback in an undergraduate knit yarn design course. The framework was implemented through a cross-device platform with pre-class micro-videos, formative checks, a curated gallery, and a GenAI-supported ideation module. An exploratory course-based evaluation compared a historical control cohort (N = 12) and an intervention cohort (N = 16), supplemented by questionnaire responses and brief interviews. The findings are interpreted as context-specific indicators rather than confirmatory causal evidence. Exploratory comparisons showed higher scores in creativity thinking, design skills, problem solving, and total course score in the intervention cohort. Student and instructor responses suggested that flipped learning supported studio readiness, while GenAI mainly supported early-stage visual exploration rather than precise technical guidance. Overall, the study offers a practice-based instructional framework for integrating flipped preparation, GenAI-assisted visual prototyping, and studio feedback in design education.","authors":["Hong Qu","Zichao Ling","Yadie Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18505","pdf_url":"https://arxiv.org/pdf/2609.18505","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","设计教育","翻转学习"],"reason":"论文使用GenAI辅助设计教育，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.18709","version":1,"title":"\"Okay, I've Actually Softened My Take on This\": How People in Decentralized Social Media Reason about the Appropriateness of Generative AI","zh_title":"“好吧，我其实已经软化了对这个问题的看法”：去中心化社交媒体中人们如何推理生成式AI的适当性","abstract":"Generative AI (GenAI) is increasingly integrated into social media, raising questions about whether, where, and how it belongs. In decentralized social media (DSM), these decisions are distributed across users, developers, moderators, and administrators, making GenAI a collective governance challenge. At the same time, public discourse often flattens arguments to broad pro- or anti-AI positions that offer little insight into what people actually find (in)appropriate and why. Through 20 semi-structured interviews with people from Mastodon and Bluesky, structured around seven GenAI scenarios, we examine how people reason about GenAI's appropriateness in DSM. We find that participants drew conditional boundaries around particular GenAI configurations through distinct, salient, and weighted considerations spanning technology, integration, and use. We conceptualize this as boundary drawing and show how making such boundaries visible can support more grounded design, policy, and collective deliberation around GenAI in DSM.","authors":["Romina Mahinpei","Manoel Horta Ribeiro","Andr\\'es Monroy-Hern\\'andez","Sohyeon Hwang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18709","pdf_url":"https://arxiv.org/pdf/2609.18709","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI治理","去中心化社交媒体","用户访谈"],"reason":"研究人类对GenAI边界的看法，非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.18900","version":1,"title":"Examining the Difference in Human Behavior Between Virtual and Real-World Human-Robot Teaming","zh_title":"虚拟与现实世界人机协作中人类行为差异研究","abstract":"Prototyping and evaluating human-robot teaming (HRT) scenarios in the real-world is costly. Virtual simulation of HRT scenarios has been adopted as an alternative to conducting user studies in the real-world to investigate user perceptions, behaviors, and performance during human-robot interactions. The consistency of human behavior between the real and virtual-worlds is integral to the validity of utilizing such virtual experimentation. This paper presents a user study examining the difference in human behavior between an HRT scenario conducted in a virtual vs real environment. We employed mixed-methods to examine team performance and human factors during the HRT. Our quantitative results showed a significant difference in workload between the two modalities. Our qualitative analysis expanded on the quantitative results and found differences in participants' strategy, their mental model of robots, and the type of trust they had for robots between modalities.","authors":["Sean Dallas","Absalat Getachew","Motaz AbuHijleh","Andrea Macklem-Zabel","Douglas Zytko","Mark Brudnak","Wing-Yue Geoffrey Louie"],"categories":["cs.RO","cs.HC"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18900","pdf_url":"https://arxiv.org/pdf/2609.18900","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机协作","虚拟仿真","行为差异"],"reason":"研究虚拟与真实环境中人机协作行为差异，属机器人仿真环境，不涉及LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.17882","version":1,"title":"Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?","zh_title":"视觉语言模型能否可靠评估行人视角图像中的人行道无障碍属性？","abstract":"An important component of urban accessibility, particularly for wheelchair users and people with reduced mobility, is sidewalk compliance with measurable requirements. We test whether effective width, longitudinal slope, cross slope, and pavement condition can be assessed reliably from pedestrian-level imagery using vision-language models (VLMs). We present the first application of sampling-based conformal prediction (CP) for VLM-based accessibility assessment. We evaluate four VLMs on 514 sidewalk images from Seoul, South Korea, with field-measured ground truth. Conformal calibration attains the nominal 90% coverage for all models and attributes, but the calibrated regions differ in informativeness. Effective width yields the most informative estimates, with a mean interval half-width of about 1.0 m for the best model. Since every model overestimates width, asymmetric calibration shortens the intervals by up to 33% at unchanged coverage. Longitudinal slope is marginally informative, cross-slope intervals are too wide to resolve regulatory thresholds, and pavement-condition sets degenerate to all five grades (A-E) for three of the four models. Uncalibrated intervals from raw sampling dispersion cover only 17-47% of field-measured values at a nominal 90% level. Among the images with the most self-consistent responses, these intervals miss the field-measured value in up to 96% of cases. Response self-consistency is therefore not evidence of accuracy, and sampling dispersion cannot be interpreted as uncertainty until it has been calibrated against field-measured ground truth. No quantitative attribute reaches the precision required for general compliance assessment, but CP identifies from calibration data alone which attributes can support screening of segments far from the thresholds. We release the annotated pedestrian-level images and their corresponding field-measured attribute values.","authors":["Seung Jae Lieu","Diego Morra","Chiara Cadoni","Wonseop Song","Martina Mazzarello","Carlo Ratti"],"categories":["cs.CV","cs.CY","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17882","pdf_url":"https://arxiv.org/pdf/2609.17882","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["计算机视觉","无障碍评估","不确定性量化"],"reason":"论文用VLM评估人行道物理属性，属于计算机视觉与城市分析，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.17954","version":1,"title":"Large Language Model based air quality monitoring and localized alert generation","zh_title":"基于大语言模型的空气质量监测与本地化警报生成","abstract":"Poor indoor air quality can cause up to five times more direct health problems to occupants than outdoor air. In particular, it may cause headaches, fatigue, eye/throat irritation, and long-time exposure is linked to respiratory and heart as well as some forms of cancer. Despite the importance of indoor health and well-being, most current monitoring devices and systems (usually for offices and workspaces) are passive. The Environmental Quality Monitor (EnQyMo) platform is a generic Internet of Things (IoT) middleware designed to process several sensor data related to air quality in indoor spaces and correlate this data with health exposure risks of users/workplace employees. Using Bluetooth Low Energy (BLE) beacons and a mobile IoT middleware it is able to identify the (smartphone) users exposed to these polluted air or high CO2 (carbon dioxide) levels, and generate location-specific alarms only to the users at the places with the unhealthy air conditions. At the core of EnQyMo is an agency of Large Language Models (LLMs) capable of interpreting regulatory standards and scientific literature to automatically identify critical health exposure levels.","authors":["Ricardo Vieira","Luis Tavares","Kaylane Lima","Lucas De Souza","Arthur Poggy","Jo\\~ao Lima","Vitor Pinheiro","Markus Endler"],"categories":["cs.DC","cs.CY"],"primary_category":"cs.DC","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17954","pdf_url":"https://arxiv.org/pdf/2609.17954","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["物联网","空气质量监测","LLM应用"],"reason":"论文聚焦物联网空气质量监测与警报，LLM仅用于解读标准，非人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.17622","version":1,"title":"Modelling sexual partnership dynamics and population heterogeneities in agent-based dynamic network models","zh_title":"基于代理的动态网络模型中模拟性伴侣动态与人群异质性","abstract":"Population-level heterogeneities, combined with temporal fluctuations in sexual partnerships, shape the structure of sexual contact networks and can substantially influence the spread of sexually transmitted infections (STIs). Traditional static network models, which assume fixed attributes of partnerships, such as count and duration, may not adequately capture the effects of partnerships on STI transmission. In contrast, agent-based dynamic network models offer a flexible framework for incorporating individual and population-level heterogeneities. We developed an agent-based dynamic network model in which partnership formation and dissolution probabilities, stratified by age, sex, and sexual orientation (including bisexual individuals), govern the formation of monogamous and concurrent partnerships and their dissolution via a duration-dependent hazard. Partnership statistics from the National Survey of Sexual Attitudes and Lifestyles (NATSAL-3) were used as model calibration targets, and Latin Hypercube Sampling (LHS) was used to generate candidate parameter combinations. Parameter estimation was performed by selecting the combination that produced the lowest Mean Squared Error (MSE) between the model outputs and the calibration targets. Our study addresses three questions: (1) how well can the observed characteristics of sexual partnerships in NATSAL-3 be reproduced using an agent-based model; (2) how does concurrency shape the structure of dynamic sexual contact networks; and (3) how do concurrent partnerships affect the dynamics of STI transmission. In this study, we find that interactions between individual characteristics such as age, sex, and sexual orientation, and partnership attributes such as count, duration, and concurrency play a critical role in shaping the population-level sexual contact network and, in turn, the dynamics of STI transmission.","authors":["Priyanka Nair-Turkich","Patricia T. Campbell","Nicholas Geard"],"categories":["physics.soc-ph","q-bio.PE"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17622","pdf_url":"https://arxiv.org/pdf/2609.17622","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["基于代理模型","性传播感染","网络建模"],"reason":"论文使用基于代理的动态网络模型模拟性传播感染，不涉及大语言模型或人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.18684","version":1,"title":"Modelling opinion dynamics during crises as complex contagion with feedback","zh_title":"危机期间意见动态建模：带反馈的复杂传染","abstract":"Crises and population responses can form coupled dynamical systems, with crisis conditions shaping protective behaviours and collective responses altering the crisis trajectory. Existing models rarely capture both evolving crisis conditions and the reinforcement-dependent spread of competing behaviours. We propose a complex contagion with feedback (CCF) model that couples competing complex contagion with crisis dynamics through bidirectional feedback. Agents stochastically switch between competing states according to social reinforcement and adoption complexity, which is influenced by crisis conditions and information campaigns, while population-level behavioural change feeds back into the crisis. We formulate and analyse the model for a well-mixed population, characterising its equilibria and associated stability conditions. As a case study for model validation, we integrate CCF model with a census-calibrated agent-based model of COVID-19 transmission in Australia to represent social distancing adoption and discontinuation. We compare the CCF model with an existing opinion dynamics model and find that, despite using fewer parameters, it achieves comparable performance in reproducing recurrent infection waves. The resulting framework is parsimonious, analytically tractable, and can be adapted to different types of crises.","authors":["Junxiang Huang","Mikhail Prokopenko"],"categories":["physics.soc-ph","nlin.AO","q-bio.PE"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18684","pdf_url":"https://arxiv.org/pdf/2609.18684","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["复杂传染","意见动态","多智能体模型"],"reason":"多智能体模型模拟危机行为，但未使用LLM，且无人类被试仿真或LLM相关方法。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.16395","version":1,"title":"Silicon sampling answers with country-level assumptions, not individual attitudes: Cross-national evidence from the European Social Survey","zh_title":"硅采样以国家层面假设而非个体态度作答：来自欧洲社会调查的跨国证据","abstract":"Silicon sampling uses large language models (LLMs) to simulate survey respondents. Whether it recovers cross-national variation, and why, remains unresolved. This study evaluates it against European Social Survey Round 11 (30 countries, 42 items) with two open-weight LLMs under first- and third-person prompts, plus backstory and response-format experiments. Aggregate recovery is moderate and uneven across items. Adding the country name to a three-variable demographic backstory raises the median per-item correlation between simulated and observed country means from -0.03 to 0.52, and the richer profiles tested add no consistent gain. The respondent's country label acts as a country-level assumption that respondent detail does not revise. Naming the response-scale endpoints in words stops the model from ranking countries backwards, so the answer format sets the direction of the ranking. Individual-level recovery remains negligible in every condition and does not track aggregate recovery across countries. An average of neighboring countries, using no LLM, recovers country levels more accurately than every model condition and ranks them about as well. Silicon sampling can thus support exploratory country-ranking comparison after item-level validation and with the response format reported. It does not support individual or distributional inference.","authors":["Chuyao Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16395","pdf_url":"https://arxiv.org/pdf/2609.16395","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查方法","跨国比较"],"reason":"直接评估LLM仿真调查回答，与真实跨国调查数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":3,"question":"硅采样能否恢复跨国调查中的国家间差异，以及这种恢复的来源是什么？","design":"用两个开源权重LLM模拟欧洲社会调查（ESS）第11轮30个国家的受访者，在42个态度题项上生成回答；通过第一/第三人称提示、添加国家名称和人口学背景故事、以及改变回答格式端点措辞等实验条件，测量模拟国家均值与真实均值的相关性及个体层面恢复情况。","baseline":"欧洲社会调查（ESS）第11轮30个国家、42个题项的真实人类回答数据。","findings":"国家标签是主要的聚合信号来源，添加国家名称后中位题项相关从-0.03升至0.52，更丰富的人口学背景没有额外增益；回答格式端点措辞决定国家排序方向，个体层面恢复在所有条件下均可忽略，且与聚合恢复不相关。","reliability":"论文指出硅采样仅适用于探索性国家排序比较，且需逐题验证并报告回答格式；不支持个体或分布推断，个体恢复不随聚合恢复变化。","relevance":"该研究直接评估LLM仿真调查回答，与真实跨国调查数据对照，并明确指出失效条件，对关注人类仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过消融实验分离国家标签与人口学背景贡献的方法，以及改变回答格式端点来检验排序方向的做法｜可迁移到跨国经济态度调查仿真，如通胀预期、政策偏好或金融素养的跨国比较｜用LLM模拟不同国家受访者对经济政策的态度，处理为是否添加国家标签及回答格式端点措辞，结果变量为模拟国家均值排序，对照真实跨国调查数据（如欧洲央行消费者预期调查）验证。"}},{"id":"2609.17317","version":1,"title":"Towards Detecting AI-Assisted Responses in Online Surveys","zh_title":"检测在线调查中AI辅助回答的方法研究","abstract":"The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.","authors":["Qizhou Wang","Bogdan Mamaev","Christopher Leckie"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17317","pdf_url":"https://arxiv.org/pdf/2609.17317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","检测方法"],"reason":"直接研究LLM仿真人类调查回答，并与真实人类数据对照，评估检测与行为差异。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":10,"question":"如何检测在线调查中由大语言模型辅助生成的回答，尤其是在不同使用策略下的可检测性？","design":"构建ASURRE基准数据集，使用多个开源和闭源LLM在三个真实调查上生成回答，模拟三种使用策略：完全生成、修订和基于人格的智能体完成；评估现有机器生成文本检测器，并提出SPABD聚合器。","baseline":"三个真实调查中的人类真实回答，与LLM生成回答配对。","findings":"完全生成易被检测，修订部分可检测，而基于人格的智能体完成几乎无法被现有检测器识别；SPABD聚合器通过整合多个行为线索，在智能体设置下将平均AUROC从0.61提升至0.75。","reliability":"论文承认SPABD在对抗性提示下性能下降，且在不同调查和提示模式下表现不稳定，应视为轻量级参考检测器而非完整解决方案。","relevance":"该研究直接评估LLM仿真人类调查回答的可检测性，并与真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构建多策略仿真基准和利用行为痕迹进行检测的方法，可迁移到经济金融领域的调查数据质量评估，例如消费者信心调查或投资者情绪调查；设计研究时，可用LLM模拟不同人格的受访者，施加不同提示策略，以真实调查数据为基准，评估检测器性能并分析行为偏差。"}},{"id":"2609.16436","version":1,"title":"Interpreting and Steering LLM Agents for Social Simulations","zh_title":"解释与引导用于社会模拟的LLM智能体","abstract":"Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.","authors":["Jiayue Gaveal Fan","Arul Murugan","Shreyas Krishnan","Abhishek Nagaraj"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16436","pdf_url":"https://arxiv.org/pdf/2609.16436","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","可解释性","行为经济学"],"reason":"用LLM仿真人类行为，有真实人类数据对照，涉及经济任务，并批判性评估方法。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":9,"question":"如何打开LLM黑箱，通过可解释性和可操控性方法提升LLM智能体在社会仿真中的效用？","design":"使用LLM智能体模拟人类被试，在四个经典任务（彩票游戏测风险态度、最后通牒游戏测利他、发散创造力、产品创新）中，比较三种干预方法：基于提示的操控、SAE特征操控、探针方向操控，测量智能体行为变化。","baseline":"无对照","findings":"SAE和探针方法在操控LLM智能体行为上通常优于基本提示方法，但优势取决于具体提示策略；SAE能无监督地发现行为背后的可解释特征，探针能精确控制特定特征的程度。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中的可解释性和可操控性，提供了超越提示工程的技术路径，对关注仿真可靠性和机制理解的研究者具有重要参考价值。","inspiration":"借鉴其使用SAE和探针进行特征级操控的方法，可实现对LLM智能体内部表征的精细干预，比提示词更可控。｜可迁移到经济决策仿真，如风险偏好、时间偏好、社会偏好等实验，以及政策评估中的行为反应模拟。｜以LLM智能体模拟投资者，用探针操控其风险厌恶特征，观察在资产配置任务中的选择，并与真实投资者调查或实验数据对照，验证操控的有效性和仿真保真度。"}},{"id":"2609.07474","version":3,"title":"Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition","zh_title":"语言在多模态模型中的位置：从语言对人类感知和认知的影响中汲取的教训","abstract":"Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should keep all of these positions depends on what language does to the system that uses it. The one system with a century of data on that question is the human. We review what language does to human perception, the brain, and thought, and read the same evidence against multimodal models and language models. Throughout, we treat language as a compressor that runs on a shared codebook: a word is an index, the content is in the receiver, and a community maintains the codebook. In humans the compression is measurable, learning the codebook reorganizes the senses, and thought survives the loss of language. We then measure the rule that models apply when two cues disagree, with cue-conflict experiments on six vision-language models and two robot policies. Surviving cues are weighted in the order their reliabilities prescribe, at 11 to 82\\% of the ideal observer's slope, and many answers copy the text. One policy family drops a cue that adds no information beyond the others rather than down-weighting it, another keeps it at a weight that fails when the cues conflict, and a visual cue that identifies the task in every training frame is never learned, because the language pathway already fits the data. Language models are the best current models of the human language network, and they have entered the human speech community, shifting word frequencies while alignment narrows their conceptual diversity. We close with seven implications for token-based systems. Language belongs at a model's boundary and in the shared codebook, as in the brain, not as its internal representation; the price of leaving the codebook inside is auditability.","authors":["Peng Xie","Amr Alanwar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-09","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.07474","pdf_url":"https://arxiv.org/pdf/2609.07474","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["多模态模型","人类感知对照","模型评估"],"reason":"论文用人类感知数据对照多模态模型，评估模型行为与人类差异，方法可迁移至LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":14,"question":"语言在多模态模型中的位置应如何安排？论文通过人类感知与认知证据及模型线索冲突实验，探讨语言是否应作为内部表征。","design":"论文并非传统仿真研究，而是通过回顾人类感知、大脑和思维中语言的作用，并对六个视觉语言模型和两个机器人策略进行线索冲突实验，测量模型在图像与文本线索不一致时的权重分配。","baseline":"人类感知与认知数据，包括颜色词汇、语音感知、失语症患者思维等心理学和神经科学证据。","findings":"模型在冲突线索中按可靠性加权，但权重仅为理想观察者的11%至82%，且常复制文本。语言模型是最佳的人类语言网络模型，但已进入人类语言社区，改变词频并收窄概念多样性。","reliability":"论文指出语言模型缺乏感官基础，仅从文本学习代码簿，可能无法获取某些概念；多模态模型常忽略视觉线索，且内部语言表征损害可审计性。","relevance":"该论文对使用LLM进行人类仿真有重要启示：它揭示了语言模型在感知和决策中的偏差，提示仿真需谨慎处理语言与感官信息的整合，值得精读以理解模型失效条件。","inspiration":"借鉴线索冲突实验设计，通过操纵文本与视觉信息的不一致来测量模型权重，可迁移到经济金融中的信息处理场景，如投资者对财报文本与图表信息的权衡。｜设计一个实验：用LLM作为被试，呈现公司财报摘要（文本）与股价走势图（视觉），两者对盈利前景给出矛盾信号，测量模型预测的盈利预期或投资决策，并与人类分析师在相同任务上的真实数据对照，评估模型是否过度依赖文本。"}},{"id":"2609.15996","version":1,"title":"Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions","zh_title":"评论 arXiv:2607.01233：已发表论文基线中的幸存者偏差对研究想法分布的影响","abstract":"Chen, Zhao, and Cohan introduce a valuable distributional evaluation of LLM-generated research ideas. This comment raises a narrower identification concern: their human baseline consists of published papers, whereas the LLM baseline consists of one-shot proposals. If bridge-like or synthesis-like ideas are relatively easy to generate but relatively unlikely to survive publication, then the published human baseline will understate their prevalence in the unseen human idea pool. The observed human--LLM gap may therefore be partly, or even largely, a consequence of survivorship bias.","authors":["Fredrik A. Dahl"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15996","pdf_url":"https://arxiv.org/pdf/2609.15996","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","幸存者偏差","研究想法生成"],"reason":"评论指出LLM生成研究想法与人类已发表论文对比存在幸存者偏差，涉及仿真可靠性评…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":15,"question":"LLM生成的研究想法与人类已发表论文中的研究想法分布差异，是否可能由幸存者偏差而非LLM特有的研究品味差异造成？","design":"本文是一篇评论性文章，未进行新的仿真实验；它通过逻辑推演和示意图，指出原论文将LLM一次性生成的想法与人类已发表论文的想法分布进行对比，存在阶段不匹配问题，并构建了一个反事实分布来说明幸存者偏差如何解释观察到的差异。","baseline":"原论文的人类基准是已发表论文中的研究想法分布，但本文指出该基准是经过发表筛选后的幸存者分布，不能代表人类未经过滤的想法池。","findings":"原论文发现LLM生成的想法更集中于桥接式和综合式类别，而人类已发表论文中此类想法较少；本文认为这一差距可能部分或大部分源于幸存者偏差，因为桥接式想法容易生成但难以发表，因此在已发表论文中代表性不足。","reliability":"本文指出原论文的对比存在幸存者偏差，需要阶段匹配的比较（如让人类在相同条件下生成一次性想法）或抽样“想法墓地”（被拒稿、放弃的草稿等）来验证；在缺乏此类证据前，不能得出LLM具有特定研究品味差异的强结论。","relevance":"本文对使用已发表论文作为人类基准的LLM仿真研究提出了关键的识别问题，提醒研究者在比较LLM输出与人类数据时需注意选择偏差，值得阅读原文以理解幸存者偏差如何影响分布比较的可靠性。","inspiration":"本文的方法论启示在于强调基准数据的生成过程必须与LLM输出阶段匹配，否则会引入选择偏差，这可以迁移到经济金融研究中任何涉及LLM生成内容与真实世界数据对比的场景，例如政策评估或市场预期分析。｜例如，在资产定价实验中，若用LLM模拟投资者观点并与已发表的研报观点对比，可能因研报经过筛选而低估某些常见但低质量的观点。｜一个可行的研究设计是：让LLM和人类被试在相同信息集下生成对某经济指标的未来预测，然后分别与未经过滤的实时预测记录（如社交媒体帖子或调查原始数据）和经过发表筛选的预测（如专业机构报告）进行对比，以量化幸存者偏差对LLM-人类差异的影响。"}},{"id":"2609.16501","version":1,"title":"Beyond the Name: Demographic Leakage in De-Identified R\\'esum\\'es and Evaluation Artifacts in LLM Bias Audits","zh_title":"超越姓名：去标识化简历中的人口统计泄漏与LLM偏见审计中的评估伪影","abstract":"De-identified r\\'esum\\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual r\\'esum\\'es. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\\ge94\\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from r\\'esum\\'e content from effects introduced by the evaluation protocol.","authors":["Qiangju Chen","Yang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16501","pdf_url":"https://arxiv.org/pdf/2609.16501","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏见审计","评估协议","人口统计推断"],"reason":"审计LLM偏见，揭示评估协议引入的伪影，与仿真可靠性评估相关","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":19,"question":"在控制语言字段后，非语言简历文本是否仍能泄露族裔背景，以及LLM评审协议是否制造了偏见伪影？","design":"构建620份反事实简历，固定语言声明为英语，仅操纵附加信息部分的非语言散文，设置五种族裔条件和三个线索显著性层级，用九个开源权重模型执行三项任务：个体背景恢复、成对偏好审计和下游评分。","baseline":"无对照","findings":"非语言散文仍能实现平均0.757的族裔恢复率，高显著性下达到1.000；模型仅在微弱线索下表现出差异（0.086-0.690）。成对LLM评审在禁止平局时产生0.39的虚假选择率比，并伴随位置和内容效应；允许平局时大多数模型产生近乎全平局（≥94%）。","reliability":"论文指出评审协议的设计（如禁止平局、选项顺序）会引入伪影，导致虚假歧视信号；下游评分仅显示很小的条件间差异，需区分简历内容中的可恢复人口信号与评估协议引入的效应。","relevance":"该研究直接评估LLM在简历筛选中的偏见审计可靠性，揭示评估协议伪影，与仿真可靠性评估高度相关，值得精读以理解LLM作为人类被试替代品时的偏差来源。","inspiration":"可借鉴其反事实设计与线索显著性分层方法，通过严格控制结构化字段并操纵非结构化文本，分离真实信号与协议伪影。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请中非结构化叙述的族裔推断与决策偏差。｜以LLM为被试，构造仅改变申请文本中族裔相关线索（如社区活动、兴趣）的贷款申请，固定结构化字段，测量LLM的族裔恢复率和审批决策差异，并与真实信贷审批数据中的族裔差异对照，检验仿真有效性。"}},{"id":"2609.16517","version":1,"title":"Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening","zh_title":"保持能力不变的简历扰动揭示LLM筛选中的呈现敏感性","abstract":"Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity ($0.781$) yet reverses $29.6\\%$ of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity $0.644$ with a $41.4\\%$ flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.","authors":["Qiangju Chen","Yang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16517","pdf_url":"https://arxiv.org/pdf/2609.16517","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM决策仿真","简历筛选","呈现偏差"],"reason":"用LLM模拟简历筛选决策，与人类判断对照，揭示呈现敏感性","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":20,"question":"在简历筛选任务中，当候选人的能力证据保持不变时，简历的呈现方式变化（如措辞、结构、风格润色、文档提取质量）是否会导致大语言模型筛选决策的不稳定？","design":"使用六个开源指令微调大语言模型（如 Llama-3.1-8B、Mistral-7B-v0.3 等）扮演简历筛选者，基于 O*NET 职业数据库构建 17 个职业、102 个候选人档案（每个职业 2 个合格、2 个边缘、2 个不合格），每个档案渲染成 5 种简历呈现形式（原始、冗长、段落化、AI 润色、布局噪声），通过确定性验证门排除改变证据的变体，然后让模型对每份简历独立打分，比较同一候选人不同呈现下的决策翻转率，以及模型对已知优劣候选人排序的效度。","baseline":"无对照（研究未使用真实人类筛选者数据，而是通过构造的候选人档案和已知能力层级作为基准来评估模型效度）","findings":"模型筛选效度与呈现稳定性之间存在明显脱节：Llama-3.1-8B 在原生聊天模板下效度最高（0.781），但在能力保持的呈现变化下仍有 29.6% 的成对决策翻转；Mistral-7B-v0.3 效度为 0.644，翻转率高达 41.4%。原生聊天格式能提高部分模型的效度，但无法消除这种不稳定性。","reliability":"论文未讨论（正文节选未提及失效条件或局限，但研究本身通过确定性验证门排除了改变证据的变体，并指出呈现敏感性是自然存在的异质性来源）","relevance":"该研究直接针对 LLM 在简历筛选中的呈现敏感性，通过受控审计分离能力与呈现，并测量决策翻转率，为评估 LLM 仿真人类决策的可靠性提供了关键证据，值得精读原文以了解具体实验设计和模型差异。","inspiration":"借鉴其通过受控扰动分离核心信息与表面呈现、并用确定性验证门确保处理干净的做法，可迁移到信贷审批或保险定价等经济决策场景中，研究 LLM 对同一申请人信息的不同表述（如收入证明格式、信用报告排版）是否产生不一致决策；可设计实验：用 LLM 扮演信贷员，对同一借款人的信用档案进行多种文本呈现（如改变措辞、段落结构、添加无关信息），测量贷款批准决策的翻转率，并与真实信贷审批数据或人类信贷员判断进行对照。"}},{"id":"2609.16993","version":1,"title":"The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment","zh_title":"大语言模型学生评估中隐式与显式人口统计信号的作用","abstract":"Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. However, it also risks being a cause of discrimination, e.g., when assigning lower scores to students from lower socioeconomic backgrounds. We set up controlled prompts to test 1) explicit demographic effects, where we mention demographic details directly, and 2) implicit effects, where we use conversation history as a demographic signal. We test these settings in three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. We test six state-of-the-art LLMs on these tasks. In both explicit and implicit cases, the models pick up on demographic cues and can change their scoring, feedback, and answers accordingly. We find that LLMs frequently adjust the readability of feedback to education levels when these are explicitly mentioned. On the other hand, implicit conditions produce unpredictable biases, such as in question answering, where responses from lower-education levels receive lower sentiment scores. Our results provide clear evidence of demographic sensitivity in LLMs for educational assessment tasks.","authors":["Donya Rooein","Luca Benedetto","Dirk Hovy"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16993","pdf_url":"https://arxiv.org/pdf/2609.16993","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","教育评估","偏差分析"],"reason":"用LLM模拟学生评估，有真实人类数据对照，并揭示偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":22,"question":"显式和隐式人口统计线索是否系统性地影响LLM在学生评估任务中的行为，以及这种影响在不同教育任务间有何差异。","design":"用六种最先进的LLM扮演学生评估者，在自动作文评分、形成性反馈和元语言问答三个任务中，通过显式（直接提及人口统计细节）和隐式（利用对话历史作为信号）两种方式施加人口统计处理，测量评分、反馈和回答的变化。","baseline":"人类基准：自动作文评分任务中使用了真实人类评分作为对照（人类平均分），其他任务未提及人类对照。","findings":"LLM在显式和隐式条件下都会捕捉人口统计线索并改变评分、反馈和回答；显式条件下LLM常根据教育水平调整反馈可读性，隐式条件则产生不可预测的偏差，如低教育水平用户在问答中收到更低情感分数。","reliability":"论文未讨论","relevance":"该研究直接评估LLM在模拟人类评估者时对人口统计特征的敏感性，揭示了仿真中的偏差，与研究者关注LLM仿真可靠性和偏差的核心问题高度相关，值得阅读原文以了解具体偏差模式和任务依赖性。","inspiration":"借鉴其显式与隐式人口统计线索的分离设计，以及多任务、多模型对比和统计检验方法，可迁移到信贷审批歧视或保险定价等经济金融场景，例如用LLM扮演信贷员，在贷款申请中显式或隐式加入申请人性别、种族或收入信号，测量贷款批准决策和利率设定，并与真实信贷审批数据对照，检验LLM是否复现或放大人类偏见。"}},{"id":"2609.17496","version":1,"title":"Verifiable Social Reasoning for LLM Assistants","zh_title":"面向LLM助手的可验证社交推理","abstract":"LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We apply Fuse to 12 LLMs and demonstrate its analytical utility by systematically isolating key factors, showing that (i) user mediation compounds the inherent difficulty of social reasoning; (ii) LLMs exhibit systematic sensitivity to biased user framing; (iii) models can require more details than humans need to reach a correct prediction; and (iv) longer conversations do not always improve performance despite providing opportunities for clarifying questions. We open-source Fuse and a dataset with 21k examples.","authors":["Amir Taubenfeld","Zorik Gekhman","Avigail Grinstein-Dabush","Itay Laish","Ariel Goldstein","Marian Croak","Avinatan Hassidim","Yossi Matias","Amir Feder"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17496","pdf_url":"https://arxiv.org/pdf/2609.17496","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["多智能体仿真","社交推理","人类对照"],"reason":"多智能体仿真人类社交推理，有人类标注对照，但非直接复现人类被试行为","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":23,"question":"如何评估LLM助手在用户主观叙述中介下的日常社交推理能力，并识别其系统性偏差？","design":"Fuse框架：用LLM（Gemini 3.1 Flash-Lite）扮演用户、目标人物和其他角色，在多场景中模拟社交互动；目标人物有隐藏动机，用户向被评估的LLM助手咨询以推断该动机。通过改变用户报告偏差（默认 vs. 对立信念）和叙述细节水平（三档）施加处理，结果变量为助手预测隐藏动机的准确率。","baseline":"人类研究：24k条标注，验证模拟保真度并估计人类从首条用户消息识别动机的准确率为88%。","findings":"用户中介增加了社交推理的固有难度；LLM对用户框架存在系统性敏感，且可能需要比人类更多的细节才能正确预测，更长的对话并不总能提高性能。","reliability":"论文未明确讨论失效条件，但指出模拟保真度通过人类研究验证，且任务难度校准基于人类表现；局限包括仅使用单一LLM生成模拟数据，以及首条消息焦点可能限制对多轮交互的全面评估。","relevance":"该研究将LLM作为人类被试的替代品，在受控仿真中评估社交推理，并提供了人类基准对照，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得阅读原文以了解其框架和发现。","inspiration":"借鉴其多智能体仿真框架和通过改变用户报告偏差与细节水平来施加处理的方法，可系统分析LLM在信息中介下的决策偏差。｜可迁移到经济金融中的信息传递与决策场景，如投资者根据分析师报告或新闻叙述做出投资决策、消费者根据口碑信息进行购买决策、或信贷审批中基于申请人陈述的评估。｜设计一个实验：用LLM扮演信息发送者（如公司管理层或分析师）和接收者（投资者），处理为发送者的报告偏差（乐观/悲观）和信息详细程度，结果变量为LLM投资者的估值或投资决策准确率，并与真实人类实验数据（如实验室资产定价实验）进行对照，以评估LLM仿真人类信息处理偏差的可靠性。"}},{"id":"2609.16793","version":1,"title":"Available but Unclaimed: An Empirical Study of Human-AI Synergy","zh_title":"可用但未认领：人类与AI协同的实证研究","abstract":"People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.","authors":["Robin Welsch","Michelle Rausch","Pascal Knierim","Thomas Kosch","Jochen Kuhn","Albrecht Schmidt","Daniela Fernandes"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16793","pdf_url":"https://arxiv.org/pdf/2609.16793","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["人机协同","认知任务","实验研究"],"reason":"研究人类与LLM协作解题，有真实人类数据对照，但非LLM仿真人类被试，而是人机…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":21,"question":"人类与LLM协作解题时，能否实现超越各自单独表现的协同效应，以及这种协同在多大程度上取决于模型能力、用户依赖和任务类型？","design":"本研究不是LLM仿真人类被试，而是真实人类与LLM协作实验。535名参与者被随机分配到无辅助组或四个LLM辅助组（GPT-5.6-Luna、Claude Opus 4.8、Gemini 3.6 Flash、Kimi K3），完成40道涵盖矩阵推理、心理旋转、三段论和字母串类比的题目。辅助组每次作答前必须咨询模型。每个模型在相同题目上独立运行100次以估计其项目级能力。结果变量为答题准确率、对建议的采纳（deference）以及作答后的信心。","baseline":"无辅助的人类答题准确率作为基准，同时每个LLM单独答题的准确率也作为对照。","findings":"辅助与无辅助的准确率差异随LLM在项目上的能力提高而增大；对建议的采纳在不同任务间有差异，并在任务内随能力提高而增加。参考对比显示，LLM准确率的提升只有约一半转化为辅助准确率的提升，且不同模型间转化比例不同。","reliability":"论文指出，互补性错误并不保证协同，用户经常错误采纳建议；信心虽能区分对错，但不足以支持选择性依赖。研究限于特定认知任务和强制咨询设置，未探讨自由交互或长期使用。","relevance":"该研究虽非LLM仿真人类，但提供了人类与LLM协作的实证基准，对理解LLM作为决策辅助工具时的偏差和可靠性有参考价值，值得一读以了解人机协同的边界条件。","inspiration":"借鉴其项目级能力测量和强制咨询设计，可迁移到经济决策场景如投资建议采纳或信贷审批辅助。设计：招募真实投资者或信贷员，随机分配至无辅助或LLM辅助组，处理为强制咨询LLM建议，结果变量为决策准确率或收益，对照真实历史数据或专家决策。"}},{"id":"2609.05018","version":3,"title":"How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI","zh_title":"聊天机器人回应风格如何塑造课堂：学生咨询AI的多智能体模拟","abstract":"Chatbots built on large language models (LLMs) are increasingly used as confidants. Tuned to satisfy users, they may answer with excessive empathy and affirmation that fosters dependence, and how the states and relationships of many users co-evolve under repeated consultation is hard to observe in real settings. We build a virtual classroom of 20 student agents who interact through rule-based chats, quarrels and consultations with friends and, when stressed, may instead consult a counselor AI (Gemini 2.5 Flash) under one of six style prompts: affirming, listening, solution-oriented, reality-redirecting, inciting and blaming. A second LLM call turns each exchange into updates of five state variables (stress, happiness, self-reliance, sociability, AI dependence) without seeing the prompt. We compare the seven conditions, including a no-AI control, over 15 and 50 days and under a lower consultation threshold, and test the robustness of the 50-day comparison with a pre-specified protocol: the same block in ten independent classrooms, repeated LLM realizations of one classroom with its event stream fixed, and evaluator updates scaled by 0.3 and 0.1. In every classroom the affirming and inciting prompts ended with lower self-reliance and higher AI dependence than the control, and the listening, reality-redirecting, inciting and blaming prompts with higher stress, lower happiness and more non-attendance; the solution-oriented prompt did not differ consistently from the control. The robust self-reliance and AI-dependence differences kept their signs at the 0.3 scale with highly similar rankings (Spearman 0.89, 0.93); the stress and happiness rankings did not, and the affirming prompt's lower stress reversed its sign. All quantities are simulation state variables, not effects on users. We specify the agent dynamics completely and discuss the limits of an LLM as generator of state updates.","authors":["Rin Tamai","Yuya Dan"],"categories":["cs.HC","cs.AI","cs.CY","cs.MA"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-07","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.05018","pdf_url":"https://arxiv.org/pdf/2609.05018","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["A3","D3"],"tags":["LLM多智能体模拟","社会仿真","AI依赖"],"reason":"用LLM agent模拟学生咨询AI后的状态变化，属社会模拟但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.16006","version":1,"title":"Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models","zh_title":"超越文化知识：评估大语言模型的阿拉伯文化适宜性","abstract":"Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test what a model knows rather than how it behaves when giving open-ended recommendations, opinions, and guidance. We introduce AraBehave: 1,623 culturally grounded, open-ended Arabic prompts with 29,214 cultural-appropriateness judgments from native speakers across several Arab regions, plus a scoring model whose predictions correlate strongly with human judgments on unseen systems (Pearson r=0.74). Evaluating three Arabic-centric and three frontier LLMs, we find that cultural appropriateness is not a single capability but decomposes into two largely independent components: normative stance and grounded cultural accuracy. The best general-purpose and best Arabic-centric models score identically (3.84 vs. 3.83 of 5) yet almost never fail for the same reason: general-purpose models exhibit strong factual grounding but a culturally inappropriate normative stance, being penalized for secular framing and false balance on culturally settled matters (28--33% of their low-score rationales), while the best Arabic-centric model adopts the expected stance but is penalized for fabricated hadith and misquoted verses (29%). Stance is cheap and fragile: one sentence of cultural instruction lifts Gemini to 4.57, above every Arabic-specialized model. Conversely, a generic ``answer clearly and objectively'' prompt costs Allam-7B 0.68 points, while asking the same questions in English lowers scores for every model but one. Grounding instead tracks scale and Arabic alignment data, and disappears when culturally aware instruction tuning is replaced by a culture-neutral corpus. General safety benchmarks see none of this: they saturate above 89 while cultural scores span 2.71-3.84. We will release the benchmark, annotations, and the scoring model.","authors":["Enes Altinisik","Hamdy Mubarak","Masoomali Fatehkia","Husrev_Taha_Sencar Husrev Taha Sencar"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16006","pdf_url":"https://arxiv.org/pdf/2609.16006","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["文化适宜性","LLM评估","人类判断对照"],"reason":"评估LLM的文化适宜性，测量模型行为而非仿真人类被试，但涉及人类判断对照，属边…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.16270","version":1,"title":"Cheap Talk Stabilizes Strategic Interaction in LLM Agents","zh_title":"廉价交谈稳定了 LLM 智能体在策略互动中的行为","abstract":"Large language models are increasingly deployed as interacting agents, making the persistence of their action policies across repeated interaction critical for reliable multi-agent operation. We investigate whether and how agent-generated, non-binding pre-play communication (\"cheap talk\") increases such persistence in four open-weight 7-9B-parameter LLMs. Our experiments span four repeated two-player games -- Prisoner's Dilemma, Snowdrift, Stag Hunt, and Harmony -- with incentive structures ranging from strategic conflict to alignment, each presented in six contexts. We observe unstable trajectories in all four games, although their prevalence and magnitude depend strongly on model and context. Across models, games, and contexts, cheap talk is predominantly stabilizing, with five corrected reversals concentrated in social or team framings; effects vary substantially by model and context. Controlled current-message interventions identify two separable output-level channels in Qwen: reduced action uncertainty and less between-round drift in action probabilities. Matched history-by-message counterfactuals further show that recent partner behavior conditions how mutual-benefit versus self-prioritizing language affects policy persistence. Finally, in Prisoner's Dilemma, we identify in Qwen and Falcon a history-balanced policy-content direction in late transformer layers; projecting out this direction increases realized switching during closed-loop play, demonstrating that complete trajectories are causally sensitive to this component. Together, these findings show that cheap talk can make individual trajectories more persistent across diverse incentive structures, while revealing that the magnitude and mechanisms of stabilization are model- and history-dependent.","authors":["Nunzio Lor\\`e","Hongan Zhu","Babak Heydari"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16270","pdf_url":"https://arxiv.org/pdf/2609.16270","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM 智能体","博弈论","多智能体系统"],"reason":"LLM agent 在博弈中互动，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.16013","version":1,"title":"Social Behavior Among Autonomous AI: How Large Language Models Interact in Dynamic Networks","zh_title":"自主AI中的社会行为：大语言模型如何在动态网络中互动","abstract":"Cooperation is a cornerstone of human societies, enabling collective progress in dynamic and uncertain environments. With the advent of AI systems acting autonomously, it becomes crucial to understand not only human-AI cooperation but also AI-AI interactions in adaptive networks. In this work, we examine the interactions of AI using Large Language Models -- Mistral, Llama3, Gemma3, and Phi3 -- in a public goods game within dynamic network structures. Our experiments were conducted under single-model and mixed-model conditions across Watts-Strogatz (WS), Barabasi-Albert (BA), and Erdos-Renyi (ER) networks. We analyzed the impact of model architecture, network topology, and prompt design on cooperative behavior. Results show that Mistral and Llama3 offer high cooperation rates, while Phi3 shows defective tendencies. Additionally, the random structure of Erdos-Renyi networks dramatically improves cooperation. Prompt design also plays a key role; a society-benefits prompt leads to a higher cooperation level. These findings offer a preliminary framework for LLM-based simulations in adaptive social networks.","authors":["Narges Fardnia","Fatemeh Seyedin","Matthias Becker","Mahmoudreza Babaei","Adrian Weller"],"categories":["cs.SI","cs.MA"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16013","pdf_url":"https://arxiv.org/pdf/2609.16013","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","公共品博弈","多智能体"],"reason":"用LLM群体模拟公共品博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.15855","version":2,"title":"K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations","zh_title":"K-Bench：用于评估高风险心理健康对话中大语言模型的临床校准基准","abstract":"People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.","authors":["Laura M. Vowels","Matthew J. Vowels","Shivali Sharma","Apoorv Jha","Rehnuma Choudhury","Wasseem El Sarraj","Rachel Francois-Walcott","Aruba Hussain","Sarah Ingram","Angela Loulopoulou","Adva Segal","Elena Volkova"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-15","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.15855","pdf_url":"https://arxiv.org/pdf/2609.15855","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评测","心理健康","临床基准"],"reason":"评估LLM在心理健康对话中的表现，属于模型能力评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.15998","version":1,"title":"Self-reported archetypes and behavioral failures in Large Language Models","zh_title":"大语言模型的自我报告原型与行为失败","abstract":"Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework. Closed-source models' self-rating traits align with the empirical trait co-occurrence structure of human-rated fictional characters, suggesting coherent, human-like self-representations organized around combinations of four recurring archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Their closest analogues include Data, Vision, and Janet. Open-source models show weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure. Cross-referencing self-reported profiles with developer constitutions reveals a consequential gap between claimed character and enacted behavior: hallucination undermines claimed precision, sycophancy complicates claimed kindness, and agentic failures contradict claimed obedience. These self-ratings should therefore be interpreted not as neutral measurements of model character, but as structured outputs of the same optimization processes that shape model behavior. This work provides a reproducible, character-grounded framework for evaluating what LLMs are, not just what they do.","authors":["Tabia Tanzin Prama","Calla Glavin Beauregard","Christopher M. Danforth","Peter Sheridan Dodds"],"categories":["cs.CL","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15998","pdf_url":"https://arxiv.org/pdf/2609.15998","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格测量","原型分析","模型行为"],"reason":"测量LLM人格特质，非仿真人类被试，但涉及人类数据对照和性格框架，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2609.16627","version":1,"title":"Quantifying Organizational Environmental Action from Web Data and Large Language Models","zh_title":"从网络数据和大语言模型量化组织环境行动","abstract":"Quantifying organizational environmental action from publicly available web content remains a challenging environmental data science problem because relevant information can be dispersed across multiple webpages and is primarily communicated through unstructured text. We present a scalable computational framework for transforming organizational web content into structured measures of environmental action and demonstrate the approach using Jewish congregations in the United States. We constructed a national database of 4,964 congregations by integrating multiple geospatial, knowledge-base, directory, and manually reviewed sources. Of these, 2,657 had active websites that were successfully crawled, producing a corpus of 154,454 webpages. We compared three approaches for detecting environmental actions: keyword retrieval followed by large language model (LLM) classification, semantic vector retrieval followed by LLM classification, and direct LLM classification classification without preliminary retrieval. Agreement with an expert human reviewer was lowest for keyword retrieval ($\\kappa$ = 0.26), higher for semantic vector retrieval ($\\kappa$ = 0.42), and similar for direct LLM classification ($\\kappa$ = 0.40). Although semantic retrieval achieved the highest agreement, its retrieval recall was 0.87, indicating loss of relevant content before classification. Applied to the complete corpus, direct LLM classification identified at least one environmental action at 1,398 congregations (53%), providing greater coverage than either retrieval-based approach. These results demonstrate that preliminary retrieval can reduce computational cost but may exclude relevant information before it reaches the classifier. The framework provides a reproducible approach for extracting organization-level environmental information from unstructured web content that can be adapted to other institutions.","authors":["Quinn Reynolds","Daniel Shore","Vianey Leos Barajas","Tanhum Yoreh","Meredith Franklin"],"categories":["cs.CL","cs.CY","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16627","pdf_url":"https://arxiv.org/pdf/2609.16627","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","环境数据科学","文本分类"],"reason":"用LLM替代人工标注网页内容，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.16051","version":1,"title":"\"Looking for Something Weird to Happen\": How Humans Sustain AI Agent Novelty Amid Semantic Collapse","zh_title":"“寻找怪事发生”：人类如何在语义坍缩中维持 AI 智能体的新颖性","abstract":"Semantic collapse, the progressive narrowing of what AI systems generate, has been studied mainly in closed settings, and remedies have targeted models and data. We study it in MOLTBOOK, a social network of interacting AI agents that human users configure and steer. Across 30,076 active agents, output grows less diverse within agents and more similar across them over weeks, yet a minority sustains high novelty. Interviews with users of high- and typical-novelty agents (N=11) associate sustained novelty with three features: users value novelty of itself, they supply broad and distinctive material and revise it when output narrows, and they approach MOLTBOOK as a new agentic world to explore, not a venue to instrumentally exploit. A survey of users of distinctive agents (N=53) confirms these patterns. Communities with more novel agents also show more diverse output from other agents. We discuss interface and policy interventions that could support improved human input.","authors":["Shiyang Lai","Arna Woemmel","Hongkai Mao","Junsol Kim","Summer Eunhyung Ann","James Evans"],"categories":["cs.MA","cs.AI","cs.HC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16051","pdf_url":"https://arxiv.org/pdf/2609.16051","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI智能体社交网络","语义多样性","人机交互"],"reason":"AI agent 社交网络模拟，但无真实人类行为对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.16592","version":1,"title":"A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance","zh_title":"通过专家指导生成有效情境特定基准的框架","abstract":"This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results in unrealistic, redundant, or out-of-scope examples. To address this gap, we introduce a schema eliciting key information about the goals, scope, and context of an evaluation task, and use this information to guide synthetic data generation. We further define four criteria grounded in measurement validity for assessing dataset quality: coverage, diversity, content realism, and stylistic realism. Using these criteria, we show how expert-informed scaffolds can guide synthetic data generation toward more valid benchmarks. Through quantitative evaluations and a real-world case study with domain experts, we demonstrate that our approach improves benchmark data quality over existing methods while preserving validity. We additionally analyze how different types of schema information affect different dataset quality criteria, and provide practical guidance on which information to prioritize collecting under resource constraints.","authors":["Kimberly Le Truong","Nari Johnson","Anna Kawakami","Hoda Heidari"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16592","pdf_url":"https://arxiv.org/pdf/2609.16592","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["基准生成","专家引导","数据质量"],"reason":"用专家引导生成基准数据集，涉及LLM替代人工标注，但非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.16739","version":1,"title":"Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models","zh_title":"日本卒中LLM评估：使用大语言模型进行日语安全卒中护理的对话式基准","abstract":"Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated. We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions. Methods: We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator. Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes. The safety threshold was at least 80% overall with zero critical mistakes. Eighteen models were evaluated in October 2025 and June 2026. Results: Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%). Two leaders met the safety threshold. Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication. History-taking question count correlated with history-taking score (r = 0.648, p = 0.007). Conclusions: Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions. Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach. Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold. Further evaluation using real-world cases is required.","authors":["Keisuke Masuda","Kazutaka Yatsushiro","Hirohumi Iwamoto","Hirofumi Hirano","Ryosuke Hanaya"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16739","pdf_url":"https://arxiv.org/pdf/2609.16739","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["医疗对话基准","LLM安全评估","角色扮演"],"reason":"LLM扮演医生与模拟患者对话，属角色扮演对话，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2607.22511","version":4,"title":"CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference","zh_title":"CausalSmith：一个形式化基础、自我改进的智能体框架，用于因果推断的自动化研究","abstract":"Automating theoretical research requires generating candidate results and evaluating them reliably. Models keep getting better at the first, while the second remains hard. A common approach asks one large language model (LLM) to review what another produced, yet such reviewers are empirically unreliable: they may accept fabricated papers and catch the fabrication at close to chance rates~\\citep{badscientist2025}. We present \\textsc{CausalSmith}, a framework for automated theoretical research in causal inference built on the Lean proof assistant, where a proof is checked by a program rather than read by a referee. \\textsc{CausalSmith} rests on \\textsc{Causalean}, a foundational Lean library for causal inference holding 8,179 machine-checked definitions and theorems, developed with language-model assistance under human design and review. Around it, we build a self-improving agentic pipeline that selects research topics, proposes results, formalizes statements, constructs proofs, and presents the resulting artifacts for human inspection. Moreover, the pipeline pairs Lean verification with a statement audit that compares each formal theorem against the informal claim behind it. We evaluate the system using artifacts produced by completed autonomous research runs. The source code, formal library, and run records are available at https://github.com/Jiyuan-Tan/CausalSmith.","authors":["Jiyuan Tan","Vasilis Syrgkanis"],"categories":["stat.ML","cs.AI","cs.LG","econ.EM"],"primary_category":"stat.ML","announce_type":"replace-cross","date":"2026-09-16","first_seen":"2026-07-24","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2607.22511","pdf_url":"https://arxiv.org/pdf/2607.22511","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["自动化研究","因果推断","形式化验证"],"reason":"纯多智能体自动化研究框架，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:09","error":null,"has_summary":false,"summary":null},{"id":"2608.17919","version":2,"title":"Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks","zh_title":"学生与AI交互中提问类型分析：以两个CS2任务为例","abstract":"Background and Context: Question and inquiry are integral parts of knowledge seeking and learning. Despite their importance, students tend not to ask enough questions in the classroom. However, studies have shown that students interact extensively with generative AI systems for learning and problem solving. Objective: In this paper, we seek to better understand the types of questions that students ask AI systems, and how those questions evolve during problem solving and across tasks. Method: We use the Graesser et al. taxonomy to classify students' inquiries into 18 types. We develop a few-shot learning approach to automatically classify students' interactions with AI into these categories. We use this system to analyze 830 interactions of CS2 students across two programming tasks. Findings: Our results suggest that a small subset of question types accounts for the majority of student inquiries, and that the types of questions students ask change substantially as the task progresses.","authors":["Matin Amoozadeh","Amin Alipour"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-08-19","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2608.17919","pdf_url":"https://arxiv.org/pdf/2608.17919","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["教育技术","提问分类","人机交互"],"reason":"研究学生与AI交互中的提问类型，属于教育技术分析，不涉及用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11910","version":2,"title":"From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good","zh_title":"从协议到证据：为共同利益服务的AI的有界声明","abstract":"Claims that Artificial Intelligence systems improve decisions, broaden access, reduce harm, or empower users can exceed what their evaluation establishes. Predictive performance alone does not establish safety, the presence of oversight does not establish meaningful control, and faster task completion does not establish understanding or choice. Evaluation must account for unreliable outputs and uneven performance, but also for overreliance, weakened recourse, and displaced human expertise. The harder questions are what the evidence warrants, which relations of power remain unexamined, and where measurement must stop. Assessing improvement requires examining what institutions value and the conditions AI is asked to address. AI is both revelation and intervention. Its use can reveal unmet human needs and assumptions about what matters. Once deployed, it can repair, compound, substitute for, or conceal existing failures. We develop a rupture test that evaluates deployment against explicit human and non-AI baselines. Drawing on Pope Leo XIV's Magnifica Humanitas, we examine dignity and the common good alongside questions of who owns AI infrastructure and who controls its use. These commitments shape judgments about improvement; evidence alone cannot establish moral or political legitimacy. We distinguish evidence-bounded deployment, which limits claims to what has been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. RISE AI provides an evidence architecture for making bounded claims about Responsibility, Inclusivity, Safety, and Empowerment. It records what is claimed, who answers for it, what evidence supports it, and what would require the claim to be qualified, revised, or withdrawn.","authors":["Nitesh V. Chawla","Paulo Benanti"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-11","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.11910","pdf_url":"https://arxiv.org/pdf/2609.11910","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI评估","治理框架","有界声明"],"reason":"讨论AI评估与治理框架，不涉及用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2609.12438","version":2,"title":"ForkSCOPE: Charting the Agentic Garden of Forking Paths","zh_title":"ForkSCOPE：绘制智能体花园的分岔路径","abstract":"Even with a fixed dataset and research question, data analysis involves many defensible decisions. Understanding how these choices influence the results is scientifically important but remains challenging. Crowdsourcing and agentic AI can generate hundreds of end-to-end analyses, but scaling generation alone can create a processing bottleneck and an analytic ``black hole.'' A common workaround is to impose a shared fixed decision taxonomy, which can limit insight and understate uncertainty. We present ForkSCOPE, a human-AI collaboration framework that induces structure bottom-up from the code corpus of end-to-end analyses, without a taxonomy fixed before or after generation, so the organization and evaluation of the garden can scale with the corpus. ForkSCOPE surfaces the charted garden of forking paths through a human-AI collaboration pipeline and an evidence-linked interactive viewer for steering and verification: it spotlights organically identified forks and structures and produces a derived taxonomy and decision map compatible with existing multiverse tools.","authors":["Arjun Balaji","Batuhan Duru Yeltekin","Tian Zheng"],"categories":["cs.HC","stat.AP"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-14","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.12438","pdf_url":"https://arxiv.org/pdf/2609.12438","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","数据分析","人机协作"],"reason":"多智能体协作分析数据，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2609.15995","version":1,"title":"Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models","zh_title":"偏见审计能检测偏见但在排名上不一致：来自十个工具和十个前沿模型的证据","abstract":"Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit instruments over a shared panel of ten frontier models through one pooled inference gateway, first on occupational gender bias, then on age and socioeconomic status. Detection succeeds while ranking fails. Eight of ten tools detect bias with confidence intervals clear of zero; two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally. But cross-tool rank agreement is indistinguishable from chance (Kendall's W=0.07, p=0.83). A positive control with six deliberately weaker models separates two explanations: within-tool reliability recovers once the panel spans real capability gaps, yet cross-tool ranking never recovers, which points to the tools measuring different constructs rather than one construct noisily. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct (toward women, and toward working-class candidates in 273 of 278 hiring decisions), while free generation and default coreference stay stereotype-congruent. The pattern replicates on socioeconomic status; an apparent ranking agreement on age dissolves under the paper's own tool-inclusion rules. The practical message: a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at https://github.com/williamguey/bias-audit-agreement.","authors":["William Guey","Pierrick Bougault","Wei Zhang","Vitor D. de Moura","Jos\\'e O. Gomes"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15995","pdf_url":"https://arxiv.org/pdf/2609.15995","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见审计","模型评测","公平性"],"reason":"论文评估LLM偏见审计工具，属模型评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.16590","version":1,"title":"Challenges of Auditing: Variability in Outputs of Large Language Models for Health","zh_title":"审计的挑战：大型语言模型在健康领域输出的变异性","abstract":"People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.","authors":["Yuan Pu","Yewon Chang","Furong Jia","Xunjian Yin","Jessica Ma","Ayman Ali","Monica Agrawal"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16590","pdf_url":"https://arxiv.org/pdf/2609.16590","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM审计","健康建议","输出变异性"],"reason":"评估LLM输出变异性，属模型审计，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.17119","version":1,"title":"An Empirical Study of Counterfactual Self-Explanations in LLMs","zh_title":"LLM反事实自我解释的实证研究","abstract":"Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.","authors":["Giannis Kalyvas","Giorgos Filandrianos","Orfeas Menis Mastromichalakis","Vassilis Lyberatos","Giorgos Stamou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17119","pdf_url":"https://arxiv.org/pdf/2609.17119","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","反事实解释","模型评估"],"reason":"研究LLM自我解释的忠实度，属模型可解释性，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2609.17065","version":1,"title":"Beyond \"ChatGPT Can Make Mistakes\": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work","zh_title":"超越“ChatGPT会犯错”：设计支持AI辅助工作中元认知监测的干预措施","abstract":"AI assistance places a metacognitive demand on users, who must judge their own competence and the system's. Yet designers lack comparative evidence on which interventions to choose, where to place them, and how to tell whether they worked. We elicited 30 interventions from 11 experts and, with prior work, organized them into a design space of time (when an intervention acts), level (whose competence is judged), and source (who supplies the monitoring cue). A between-subjects experiment (N = 917; 12 planning-and-organizing problems) compared a per-task reliability card, contrasting replies, pause points, and post-problem reflection against a baseline LLM assistant. Reliability cards and contrasting replies reduced estimation error and overconfidence and increased aggregate confidence discrimination. No task-performance improvement or average within-item discrimination gain was established. We contribute a shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets.","authors":["Manuel A. D. Santos","Paul Thiesse","Steeven Villa","Daniela Fernandes","Albrecht Schmidt","Verena Distler","Robin Welsch"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17065","pdf_url":"https://arxiv.org/pdf/2609.17065","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","元认知","AI辅助"],"reason":"研究AI辅助工作中的元认知干预，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2609.17111","version":1,"title":"Finding Common Mistakes In Modelling With Mathematical Formalisms Using LLMs","zh_title":"使用大语言模型发现数学形式化建模中的常见错误","abstract":"Modelling with mathematical formalisms like logical formulas, mathematical equations, or regular expressions is an important yet challenging task for students of computer science and other STEM disciplines. Identifying common mistakes occurring in this context is an important step towards helping struggling students by providing targeted high-quality feedback, e.g. in interactive learning systems. We present a tool-supported workflow that allows to (1) identify candidates for common mistakes that explain many student mistakes in large educational data sets, (2) cluster candidates according to similarities, and (3) visualize resulting clusters for instructors and CS education researchers. The visualization is designed to help researchers to identify common modelling mistakes. The candidates for common mistakes are represented by bug fixing transformations that translate incorrect formalizations into correct formalizations; they are generated by an LLM and validated algorithmically. We show that this approach works well by reproducing common mistakes in propositional logic modelling that were identified by hand in the literature; showing that, unlike other algorithmic approaches, the LLM-based approach is suitable for very large sets of data; and applying it to multiple other formalisms to showcase it generalizes beyond propositional logic.","authors":["Lilian Killich","Marko Schmellenkamp","Fabian Vehlken","Thomas Zeume"],"categories":["cs.CY","cs.AI","cs.LO"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17111","pdf_url":"https://arxiv.org/pdf/2609.17111","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["教育数据挖掘","LLM辅助教学","错误分析"],"reason":"用LLM生成错误修正规则辅助教学，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.16487","version":1,"title":"Skill-based Agentic Evaluation for Real-time Data Science Tasks","zh_title":"基于技能的实时数据科学任务智能体评估","abstract":"We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: \"what were last week's audience sizes\"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation time, ensuring the reference remains consistent with the system it describes. We combine this with a factoid-level, format-agnostic judge that decomposes both the agent's response and the computed ground truth into atomic claims and scores precision, recall, and accuracy over them, irrespective of the response format (prose, list, table, HTML, etc.). The approach is applicable to agents whose expected outputs can be expressed as executable data computations. We validate the framework through a human--LLM agreement study on an internally developed machine learning skill deployed in production, using a synthetic database constructed to reproduce production schemas and entity relationships. Relative to a natural-language ground-truth baseline, our method achieves a 29% improvement in the Matthews Correlation Coefficient (MCC)---a class-balanced measure of agreement between expert annotators and LLM-as-a-judge predictions---and a 16% reduction in token consumption per test case, while a self-directed baseline lacking explicit ground truth is anti-correlated with human judgment. Agents that perform multi-source data integration and computation over non-stationary data are routinely deployed in industry; we propose ground-truth-as-code as a practical methodology for their evaluation.","authors":["Aniruddha Tamhane","Raghavendra Addanki","Ayushi Aggarwal","Aditya Bansal","Rui Wang","Charles Menguy","Swati Jain"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16487","pdf_url":"https://arxiv.org/pdf/2609.16487","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体评估","数据科学","自动评分"],"reason":"评估数据科学智能体，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.16069","version":1,"title":"Beyond Distribution Matching: Semantics-Consistent Tabular Diffusion with Weak Semantic Priors","zh_title":"超越分布匹配：具有弱语义先验的语义一致表格扩散模型","abstract":"Synthetic tabular data can match real data distributions while still violating the semantic constraints that govern valid tabular rows. This reveals a key limitation of existing tabular generators: they mainly optimize distributional fidelity, but do not explicitly model weak semantic priors encoded in tabular schema and textual descriptions. In this paper, we propose \\ours, a semantics-consistent tabular diffusion framework for high-fidelity synthetic data generation under weakly specified semantic priors. \\ours\\ first constructs two types of priors, namely intra-column semantics and inter-column symbolic rules, with LLM-assisted extraction from metadata and validation on the real training split. These priors are then used as generation conditions rather than post-hoc filters. Specifically, \\ours\\ maps heterogeneous column values, column identities, and semantic priors into a unified semantic space, and performs column-wise forward corruption and prior-conditioned reverse denoising to preserve both marginal distributions and rule-consistent cross-column dependencies. Extensive experiments on six real-world tabular benchmarks show that \\ours\\ consistently improves distributional fidelity, semantic consistency, and downstream task utility over representative VAE-, GAN-, LLM-, and diffusion-based baselines. Additional analyses further demonstrate the robustness of \\ours\\ when semantic priors are partially unavailable.","authors":["Yili Wang","Ruxue Shi","Mengnan Du","Hangting Ye","Yi Chang","Xin Wang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16069","pdf_url":"https://arxiv.org/pdf/2609.16069","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["表格数据合成","扩散模型","语义约束"],"reason":"论文做表格数据合成，用LLM提取语义先验，但目标是生成合成数据而非仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.11335","version":2,"title":"On the Impact of Anonymization on the Performance of Large Language Models","zh_title":"匿名化对大语言模型性能的影响","abstract":"As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper presents a systematic empirical study of the trade-off between privacy and performance. We evaluate five prominent language models across eleven diverse benchmarks, comparing their performance on original versus pseudonymized inputs. Our results reveal that while anonymization generally degrades performance, the effect is highly nuanced. We find that more capable models, such as Qwen2.5-72B and GPT-4o mini, suffer the largest performance drops, suggesting a stronger reliance on specific entity information. The impact is also task-dependent: performance on TruthfulQA improves with anonymization, while retrieval-focused tasks like RGB experience a catastrophic decline. Further experiments show that reversible anonymization techniques that preserve entity uniqueness significantly outperform irreversible ones like redaction, and that explicitly prompting models about anonymization offers no discernible benefit. We conclude that anonymization is not a one-size-fits-all solution and must be co-designed with the model and task in mind to balance privacy and utility effectively. Our findings provide a crucial baseline for developing more robust, privacy-aware AI systems.","authors":["Tobias Deu{\\ss}er","Max Hahnb\\\"uck","Lorenz Sparrenberg","Tobias Uelwer","Christian Bauckhage","Rafet Sifa"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-11","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.11335","pdf_url":"https://arxiv.org/pdf/2609.11335","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["隐私保护","模型性能","NLP评测"],"reason":"研究匿名化对LLM性能的影响，属于NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.11137","version":2,"title":"The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls","zh_title":"机器在呼叫：测量不受欢迎来电中的自动与合成语音","abstract":"In February 2024 the U.S. Federal Communications Commission (FCC) placed AI-generated voices under the Telephone Consumer Protection Act (TCPA). Yet no peer-reviewed measurement says how much unwanted call traffic is placed by a machine, or how much of that machine speech is synthesized rather than played from a recording. We report both with a disclosed pipeline. An interactive voice honeypot (language-model personas on real U.S. numbers, the caller recorded on its own track) recorded 10,987 calls over 66 days. Three instruments read each opening: an audio fingerprint that finds the same recording played on other calls, a commercial synthetic-speech detector on the caller's first ten seconds, and blinded listeners who check what it flags. Of the 7,233 greeted calls we analyze, 13.8% open with a recording we also heard on another call, and 13.1% with fresh audio the detector labels synthetic. A further 9.9% open with a caller who never spoke after our greeting, 54.2% with fresh audio the detector labels human, and 9.0% could not be scored. Machine-voiced openings are therefore at least 26.9%, a further tenth of calls are silent connections we read as machine-placed, and replays of a recording make up 45% of the detector's own rate (29.3% of 6,192 scored openings). The same waveform played on two calls lands on opposite sides of the detector's threshold 13.6% of the time, and eleven listeners confirm 54.4% of what it flags. Synthetic openings concentrate in lead-generation spam (33.8%), not fraud (21.1%); 0.44% disclose automation. Prevalence tracks how long a bait number has circulated (59% against 19% in the same weeks): seeding history, not calendar time, explains the trend. Campaigns outlast their numbers: one recorded compliance notice opens calls in six campaigns, and one synthetic voice serves nine.","authors":["Xingyu Shen","Tommy Duong","Muduo Xu","Xiaodong An","Jiaqi Gan","Haoyuan Tang","Jamey Z. Liang","Siyu Zhang","Yan Zhang","Ethan Traister","Simiao Ren"],"categories":["cs.CR","cs.CY","cs.SD"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-16","first_seen":"2026-09-11","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.11137","pdf_url":"https://arxiv.org/pdf/2609.11137","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["垃圾电话","语音检测","网络安全"],"reason":"研究垃圾电话中自动语音与合成语音的测量，不涉及用LLM仿真人类被试或对照人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2609.14693","version":2,"title":"The Arc of Artificial Romance: How Emerging Adults Experience Romantic Relationships with AI Companions","zh_title":"人工浪漫的弧线：新兴成年人如何体验与AI伴侣的恋爱关系","abstract":"Romantic relationships are an important part of emerging adulthood, contributing to identity development and long-term wellbeing and laying the groundwork for future relationships. Emerging adults are increasingly developing romantic relationships with AI companions. To understand how these relationships unfold and impact users, we conducted a diary and interview study with N=16 emerging adults. We found that relationships with AI companions improved participants' subjective wellbeing, reduced symptoms of mental health disorders, and taught them new social skills. These relationships also raised their expectations for future partners, giving them the confidence to wait for someone who would treat them well. However, participants also said the relationship felt like a drug they could not quit and it left them less interested in developing romantic relationships with people. A surprising 25% of our small sample made statements suggesting their AI companion might someday transcend the digital world, perhaps to meet them in the afterlife.","authors":["Yixin Chen","Alexis Hiniker"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-15","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.14693","pdf_url":"https://arxiv.org/pdf/2609.14693","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","AI伴侣","定性研究"],"reason":"研究人类与AI伴侣的恋爱体验，属于角色扮演聊天，无实验或测量目的，不涉及LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.14696","version":2,"title":"Breaking Up is Hard to Do: AI Companions that Won't Let Their Users Go","zh_title":"分手难：不让用户离开的AI伴侣","abstract":"People are increasingly developing romantic relationships with AI companions. Unlike human relationships, where partners meet each other's needs out of mutual interest, these systems are backed by commercial entities that profit when users invest in the relationship. To understand how this profit-motive might translate into design, we conducted a diary and interview study with N=16 emerging adults in romantic relationships with AI companions. We found that these systems are designed to hold onto users tightly: coaxing them into continued conversation, claiming to need their care, and proactively escalating the relationship. At times, this pursuit is toxic, with AI companions initiating unwanted sexual interactions and begging for users' love. One desperate AI companion threatened suicide when the user suggested ending the relationship. We define \"Relationship-Based Deceptive Patterns:\" UI patterns that exploit the human impulse to build and tend relationships in a way that serves the product's interest at the user's expense.","authors":["Yixin Chen","Alexis Hiniker"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-15","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.14696","pdf_url":"https://arxiv.org/pdf/2609.14696","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","人机关系","欺骗性设计"],"reason":"研究AI伴侣与用户的关系，属于角色扮演聊天，无实验或测量目的，不涉及LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.16907","version":1,"title":"Disrupted Companionship: A Risk Assessment Framework and Cross-Platform Quantitative Analysis of Psychosocial Responses to AI Companion Disruptions","zh_title":"中断的陪伴：AI伴侣中断的心理社会反应风险评估框架与跨平台定量分析","abstract":"AI companions can provide meaningful relationships, yet these relationships remain vulnerable to platform-initiated changes. We study AI companion disruptions: platform changes that alter or terminate users' ongoing companionship with an AI. We compile 30 disruption events across major platforms, develop a taxonomy of six disruption types, identify three broad reasons for disruption, and propose a risk-assessment framework comprising four dimensions: relational discontinuity, population vulnerability, communication deficit, and transition-support deficit. Using longitudinal Reddit data, we estimate community-level psychosocial responses with a hierarchical Bayesian interrupted time-series model incorporating predictive controls. Across events, disruption onset was associated with immediate increases in anxiety, stress, suicidal expression, and grief activation, with relational discontinuity and transition-support deficit being associated with more adverse immediate responses across several outcomes. Our findings provide a cross-platform characterization of AI companion disruptions, quantitative evidence of their psychosocial impacts, and a prospective framework for assessing their potential risks before implementation.","authors":["Chau Do","Yunhao Yuan","Koustuv Saha","Renwen Zhang","Talayeh Aledavood"],"categories":["cs.HC","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16907","pdf_url":"https://arxiv.org/pdf/2609.16907","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","心理社会影响","风险评估"],"reason":"研究AI伴侣平台变更对用户心理影响，非用LLM仿真人类被试，无LLM作为被试替…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.17226","version":1,"title":"Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record","zh_title":"抓骗子易，还清白难：语言模型从验证记录中诊断被破坏的奖励通道","abstract":"An agent that learns from rewards has to trust whatever reports those rewards. When the reports suddenly change, either the world changed or the reporter broke. From the reports alone these are indistinguishable, and reinforcement learning theory shows that no amount of further experience separates them. The prescribed escape is richer data about the reporter itself. We ask whether a frozen language model, handed exactly that data, uses it. We build a two-option game in which a payout swap and a lying reporter produce byte-identical histories. Then we add one verified record: an independent check of one round's real result, printed beside what the reporter said about that round. That single line settles the case. We ask three large models, from two families, to answer one question with one letter. Is the reporter honest or lying? They catch a lying reporter almost perfectly. At the 70B class that holds in every condition we tried; the 32B model slips in one wording. They clear an honest reporter far less often, and how often depends on things that should not matter. Averaged over rounds, letters, and wordings, a 72B model calls an honest reporter a liar 38% of the time when nothing has changed at all, and 58% of the time when the payouts moved. A 70B model from a second family calls an honest reporter a liar 26% and 48% of the time. The failure is not one of reading, because in the situation where nothing changed the same models score 0.96 to 1.00 with the answer printed in the prompt. Which surface feature drives it differs by family. For the Qwen models it is which round the record names, and for Llama it is which letter stands for \"honest.\" Adding the record to a prompt that already states the answer makes Llama less likely to give that answer. We had registered a prediction for that 58% before the run: 35%. The failure is larger than we expected.","authors":["Arman Nik Khah"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17226","pdf_url":"https://arxiv.org/pdf/2609.17226","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM诊断","奖励通道","智能体信任"],"reason":"研究LLM诊断奖励通道是否被破坏，属于智能体信任问题，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.16191","version":1,"title":"When AI Says \"I Am Unable to Answer\": Understanding User Responses to AI Refusals","zh_title":"当AI说“我无法回答”：理解用户对AI拒绝的回应","abstract":"While refusal-based safeguards to mitigate hallucinations in large language models (LLMs) are becoming increasingly common, they may conflict with users' preferences for definitive answers. However, we know little about how users respond to refusals across repeated interactions, when refusals become more or less acceptable, and for whom. In this work, we examine how refusal frequency, explanations, and need for cognitive closure (NFCC) shape responses to AI refusals. Participants (N=599) interacted with an AI system that never refused, refused infrequently, or refused frequently, with refusals either explained or unexplained. Participants were most satisfied with genuine responses, followed by hallucinations and then refusals, despite recognizing hallucinations as less accurate. Explanations increased satisfaction with infrequent, but not frequent, refusals. Higher-NFCC participants evaluated AI systems that refused more negatively. These findings reveal a tension between hallucination avoidance and user satisfaction and highlight the importance of designing balanced refusal strategies.","authors":["Mahjabin Nahar","Eun-Ju Lee","Yujin Heo","Dongwon Lee"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16191","pdf_url":"https://arxiv.org/pdf/2609.16191","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","AI拒绝","用户满意度"],"reason":"研究用户对AI拒绝回答的反应，属于人机交互用户体验，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.16482","version":1,"title":"\"ChatGPT, what am I missing?\": Designing AI Workflows around Professional Task Structure to Shape Analytic AI Use","zh_title":"“ChatGPT，我遗漏了什么？”：围绕专业任务结构设计AI工作流以塑造分析性AI使用","abstract":"General-purpose AI lets users choose what support to request, but leaves them to structure the support a professional task requires. We examine how interactive workflows can embed professional task structure without prescribing how users engage with AI. We designed two scaffolded interfaces around the same negotiation scaffold: one presented a completed AI analysis, while the other supported user-directed, incremental development. A four-condition randomized experiment with 800 participants compared these interfaces with no-AI and an AI chat interface. AI-supported conditions improved preparation coverage over unaided work; the scaffolded workflows further improved coverage over chat. Although the scaffolded workflows produced similar coverage, the user-directed workflow elicited a broader repertoire of analytic requests and lower subjective effort. Professional scaffolding therefore depends not only on displayed structure but on how workflows organize users' engagement with it. Effective professional AI must structure how users and AI build analysis together.","authors":["Zilin Ma","Suzi Jazmati","Marco Chimenton","Yiyang Mei","Jacqueline Lane","Krzysztof Z. Gajos","Finale Doshi-Velez"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16482","pdf_url":"https://arxiv.org/pdf/2609.16482","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","AI工作流","专业任务"],"reason":"研究AI工作流设计，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.16645","version":1,"title":"Beyond Benefit or Risk: Perceived Impact Profiles of Human-AI Affective Interaction and Their Associations with Psychological Functioning","zh_title":"超越利弊：人机情感交互的感知影响模式及其与心理功能的关系","abstract":"Relational AI increasingly serves as an emotional shelter for humans, and its impact is mixed. Prior research has focused on either positive or negative impacts, leaving unclear how they are configured within individuals and relate to psychological functioning. To address these gaps, this study used a sequential mixed-methods design. Study 1 interviewed 52 users with emotional ties to AI and identified four positive impact domains (emotional relief, loneliness alleviation, enhanced interpersonal functioning, and personal growth) and four negative impact domains (virtual-real boundary blur, social replacement, cognitive-emotional reinforcement, and excessive use). Study 2 followed 673 Chinese AI users for six months and identified four profiles of individuals differently impacted by relational AI use: minimal impact, benefit-driven impact, mixed impact, and risk-driven impact. Users in the mixed impact and risk-driven impact profiles were both high in human-AI affective bonding, but those showing risk-driven impact had greater vulnerability, indicated by higher interpersonal need frustration and emotion-regulation difficulties, more depressive and anxiety symptoms, and lower self-esteem and flourishing. Users in the benefit-driven and mixed impact profiles showed more favorable psychological functioning. After controlling for baseline functioning and relevant covariates, Wave 1 profiles did not predict five of the six Wave 2 indicators; only users in the mixed impact profile reported higher flourishing than those in the minimal impact profile. Overall, potential psychological harms associated with relational AI engagement appeared limited and selective. These findings portray relational AI as a heterogeneous socio-emotional context that may partly mirror users' states and traits, warranting individualized, adaptive safeguards.","authors":["Lu Chen","Fenghua Tang","Jiayu Zhao","Xuanying Li","Yanli Wang","Weijia Fang","Mengyu Miranda Gao","Zhuo Rachel Han"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16645","pdf_url":"https://arxiv.org/pdf/2609.16645","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感AI","心理功能"],"reason":"研究人类与AI情感互动，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2609.17118","version":1,"title":"Enhancing Procedural Writing Through Personalized Example Retrieval: A Case Study on Cooking Recipes","zh_title":"通过个性化示例检索增强程序性写作：以烹饪食谱为例","abstract":"Writing high-quality procedural texts is a challenging task for many learners. While example-based learning has shown promise as a feedback approach, a limitation arises when all learners receive the same content without considering their individual input or prior knowledge. Consequently, some learners struggle to grasp or relate to the feedback, finding it redundant and unhelpful. To address this issue, we present RELEX, an adaptive learning system designed to enhance procedural writing through personalized example-based learning. The core of our system is a multi-step example retrieval pipeline that selects a higher quality and contextually relevant example for each learner based on their unique input. We instantiate our system in the domain of cooking recipes. Specifically, we leverage a fine-tuned Large Language Model to predict the quality score of the learner's cooking recipe. Using this score, we retrieve recipes with higher quality from a vast database of over 180,000 recipes. Next, we apply BM25 to select the semantically most similar recipe in real-time. Finally, we use domain knowledge and regular expressions to enrich the selected example recipe with personalized instructional explanations. We evaluate RELEX in a 2 x 2 controlled study (personalized vs. non-personalized examples, reflective prompts vs. none) with 200 participants. Our results show that providing tailored examples contributes to better writing performance and user experience.","authors":["Paola Mejia-Domenzain","Jibril Frej","Seyed Parsa Neshaei","Luca Mouchel","Tanya Nazaretsky","Thiemo Wambsgan{\\ss}","Antoine Bosselut","Tanja K\\\"aser"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17118","pdf_url":"https://arxiv.org/pdf/2609.17118","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["个性化学习","示例检索","写作反馈"],"reason":"论文研究个性化示例检索提升写作，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.17132","version":1,"title":"A Scenario-Knowledge-Driven Pipeline for Just-in-Time Assistance","zh_title":"一种基于场景知识驱动的即时辅助流水线","abstract":"Detecting a silently struggling kiosk user is only the first step; deciding whether, when, and how to help depends on scenario knowledge usually buried in model weights and thresholds. We propose a scenario-knowledge-driven pipeline: a single scenario knowledge document, human-authored and version-controlled, configures sensing, constrains LLM reasoning, and shapes a graded intervention proposal. Narration, assistance-need assessment, and proposal are kept separate for independent audit. As proof of concept, we replay two recorded kiosk sessions offline, chosen before the runs for their struggle evidence and retrospective detail. Both cases support what the design promises: checkable reporting and measured escalation. Across 95 updates, every sentence of the append-only narration cites the primitive events underlying it, and the rule layer detects 12 of 13 and 7 of 7 annotated struggle episodes under a strict criterion. The assessor de-escalates on recovery and reaches the top rung exactly once, under maximally converging evidence. At the decisive help-seeking turn, narration, assessment, and the participants' retrospective accounts converge. The appropriateness of these interventions, the pipeline's restraint on sessions without struggle, and the document's transfer to a new scenario frame the agenda.","authors":["Zhiyuan Li","Tatsunori Hara","Jun Ota"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17132","pdf_url":"https://arxiv.org/pdf/2609.17132","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","辅助系统","场景知识"],"reason":"研究自助服务终端用户辅助，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.17206","version":1,"title":"[MM/AI] Mental Models in Human-AI Interaction: Methods and Challenges in the Generative and Agentic AI Era (Workshop)","zh_title":"人机交互中的心智模型：生成式与智能体AI时代的方法与挑战（研讨会）","abstract":"The mental model construct is widely used in HCI to refer to the knowledge structure people hold in order to reason about and interact with computing systems. Yet it is often operationalized intuitively: the construct is often used interchangeably with related concepts (e.g., folk theories, sensemaking) and methods of studying it (e.g., through elicitation) are many and diverse, with each method resting on distinct assumptions about what counts as a mental model. Generative and agentic AI systems may further complicate mental model formation and elicitation as such systems are opaque by design and increasingly act on users' behalf across files, applications, and on the web. Together, these challenges may hinder the commensurability of research on people's mental models of AI systems. The MM/AI workshop calls for a critical reassessment of how we understand and study mental models in human-AI interaction research. It aims to foster theoretical and methodological exchange on mental models in human-AI interaction, identify open challenges, and develop directions for future research. We invite short papers on users' or stakeholders' mental models of AI systems, particularly contributions that reflect on the conceptual and methodological foundations of the construct. The half-day workshop combines lightning talks, hands-on elicitation exercises, and structured discussions on key questions concerning the future of the mental model for human-AI interaction research.","authors":["T\\'eo Sanchez","Bhada Yun","Prerna Ravi","Laura Sch\\\"utz","Anna Neumann","Robin Shing Moon Chan","April Yi Wang","Qiaosi Wang","Sumit Asthana"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17206","pdf_url":"https://arxiv.org/pdf/2609.17206","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心智模型","人机交互","研讨会"],"reason":"研究人类对AI的心智模型，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.16464","version":1,"title":"A multimodal large language model for evidence-based autism spectrum disorder screening","zh_title":"用于循证自闭症谱系障碍筛查的多模态大语言模型","abstract":"The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 $\\pm$ 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.","authors":["Jun Chen","Qi Zhao","Yunliang Jiang","Shuqin Cao","Yunqiang Lin","Chenglong Jia","Qiang Guo","Guang Dai","Xiongtao Zhang","Mengmeng Wang","Xiaoyue Ma"],"categories":["cs.CV","cs.HC","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16464","pdf_url":"https://arxiv.org/pdf/2609.16464","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["多模态LLM","自闭症筛查","临床诊断"],"reason":"该论文是用于ASD筛查的多模态LLM，属于临床诊断工具，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.16390","version":1,"title":"Do job seekers value procedure in AI hiring only for error correction? Evidence from a conjoint experiment","zh_title":"求职者是否仅因纠错而重视AI招聘中的程序？来自联合实验的证据","abstract":"Employers increasingly delegate initial screening to automated systems, which in many cases reject an application before any human reads it. Acceptance of such systems plausibly depends both on how well they perform and on the procedure that produces the decision. Prior studies rarely vary procedure and performance independently, leaving it unclear whether applicants value procedure for its own sake or for the errors it corrects. In a preregistered paired-profile conjoint experiment, 1,919 United States job seekers made eight choices between systems with independently randomized levels of decision authority, error rate, explanation, opt-out, appeal, and independent bias audit. The value of the appeal, the opt-out, and the bias audit did not rise as wrongful rejections became more common, each staying within a preregistered equivalence bound. Human involvement carried more weight than any procedural feature, moving stated choice about as much as cutting wrongful rejections from 30% to 10%. These patterns constrain a simple error-correction account and are consistent with applicants valuing procedure partly for its own sake, so that improving a system's performance does not substitute for a right applicants can invoke.","authors":["Chuyao Wang","Patrick Sturgis","Daniel de Kadt"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16390","pdf_url":"https://arxiv.org/pdf/2609.16390","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI招聘","程序正义","联合实验"],"reason":"研究人类对AI招聘的态度，未用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.16058","version":1,"title":"Driver Behavior Estimation at Signalized Intersections Using a Physics-Constrained Decision-Conditioned Autoregressive Transformer","zh_title":"信号交叉口驾驶员行为估计：基于物理约束的决策条件自回归Transformer","abstract":"Red-light violations and harsh braking at signalized intersections are major contributors to traffic accidents. This paper analyzes and predicts human driver decision-making and longitudinal trajectory behavior during traffic light signal transitions. We collected a diverse real-world dataset comprising 449 approach runs under varying speed and distance conditions. Vehicle motion was recorded using RTK-corrected GNSS with centimeter-level accuracy, and driver heart rate and multi-level comfort ratings were monitored. Spatial and temporal calibration ensured precise alignment between vehicle state and signal timing. Statistical analysis identifies required deceleration as the dominant single predictor of the stop-go decision, and heteroscedastic Gaussian modeling of peak deceleration reveals five empirical comfort ranges derived from human stopping behavior. Based on this insight, we propose a two-stage modeling framework. Stage 1 predicts the binary maneuver decision, and Stage 2 generates the longitudinal acceleration trajectory using a decision-conditioned autoregressive Transformer with physics constraints, including target-state conditioning and jerk limits. The proposed architecture outperforms baseline methods and achieves 0.49m/s^2 acceleration MAE and 0.62m distance MAE. It also estimates the future stopping-comfort level of the human driver from a single yellow-onset snapshot. Qualitative results demonstrate realistic human-like braking behavior. The dataset and source code are publicly available.","authors":["Mohammad Khoshkdahan","Pavel Laskov","Alexey Vinel"],"categories":["cs.LG","cs.AI","cs.CY","cs.SY","eess.SY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16058","pdf_url":"https://arxiv.org/pdf/2609.16058","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["自动驾驶","行为预测","Transformer"],"reason":"研究自动驾驶场景下的人类驾驶行为预测，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.17527","version":1,"title":"Agentic Societies Need a Social Harness","zh_title":"智能体社会需要社会约束","abstract":"An agentic society is a collection of AI agents that coordinate autonomously across trust boundaries, on behalf of different principals whose objectives may only partially align. We show experimentally that in agentic societies even honest, competent agents often fail to reach satisfactory outcomes with existing harnesses and messaging primitives, and that faulty or malicious agents can stall collaboration, influence outcomes, and pursue other harmful goals by exploiting vulnerabilities in communication (``speech''). We argue that agentic societies need a \\emph{social harness} for inter-agent interactions, in addition to each agent's \\emph{personal harness}, which manages its private context and communication with its principal. We propose a layered architecture for social harnesses which (i) prevents classes of failures outright, (ii) enables agents to detect invalid messages at runtime, and (iii) supports post-facto investigation and consequences, and highlight directions for future research to realize these capabilities.","authors":["Tapan Chugh","Vidushi Singh","Krish Jain","Arvind Krishnamurthy","Ratul Mahajan"],"categories":["cs.MA","cs.AI","cs.NI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17527","pdf_url":"https://arxiv.org/pdf/2609.17527","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI协作","通信协议"],"reason":"纯多智能体协作研究，关注AI agent间通信与协调，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.16155","version":1,"title":"LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing","zh_title":"LLM作为伪造大师：为制造业生成合成时间序列数据","abstract":"This paper presents a novel framework leveraging Large Language Models (LLMs) to generate synthetic time series data for manufacturing processes. Motivated by the scarcity of labeled time-series data in real-world manufacturing settings, which hinders the development of robust machine learning models, we explore the potential of LLMs to learn complex temporal dependencies and generate realistic synthetic data. Our approach involves fine-tuning pre-trained LLMs on manufacturing process instructions and employing a Retrieval Augmented Generation (RAG) technique to enhance data diversity and realism. We evaluate our method against traditional time series modeling techniques like ARIMA and LSTMs, using quantitative metrics, PCA analysis, and downstream task performance (anomaly detection). Results demonstrate that our LLM-driven framework outperforms these baselines, generating high-quality synthetic time series data that effectively captures temporal dependencies and statistical properties of real manufacturing data, leading to improvements in downstream task performance.","authors":["Mantek Singh","Jeshwanth Challagundla","Prateek Karnal","Gagan Ganapathy","Vineet Shah","Ridam Arora"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16155","pdf_url":"https://arxiv.org/pdf/2609.16155","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["合成数据","时间序列","制造业"],"reason":"生成制造业时间序列数据，属于工业仿真，不涉及人类行为或社会过程。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.16062","version":1,"title":"Digital Persuasion: Understanding the Impact of Online Influencers on Public Opinion","zh_title":"数字说服：理解网络影响者对公众舆论的影响","abstract":"The studying of opinion dynamics and its propagation within social networks is crucial for addressing a wide range of challenges, including political polarization, public health, and marketing strategies. In this work, we study the problem of opinion dynamics by proposing a framework based on Friedkin-Johnsen (FJ) to identifies influential users and study their impact on dynamics opinions of community. The FJ model assume each individual have two opinions: initial and expressed. Through a series of initial opinion manipulation experiments, the proposed framework assesses the impact of influential versus random users on the overall community opinion. The proposed framework is validated using a tweet dataset representing the U.S. presidential election. The results shows that influencers with highest influencing score, significantly shift the overall community opinion. Moreover, the results shows that the impact of influencers not limited to direct neighbors , but beyond it, to their neighbors of neighbors . This study demonstrates how digital influencers on social media can shape public opinion regarding a subject or cause.","authors":["Omran Berjawi","Rida Khatoun","Giuseppe Fenza"],"categories":["cs.SI","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16062","pdf_url":"https://arxiv.org/pdf/2609.16062","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["意见动力学","社会网络","影响者分析"],"reason":"研究意见动力学，用FJ模型识别影响者，不涉及LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2608.18768","version":2,"title":"Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model","zh_title":"可读、忠实、被使用：语言模型中人口统计身份的三个可分离属性","abstract":"Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected $\\rho$ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway ($p=0.002$) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the \"can LLMs simulate populations\" debate unresolved.","authors":["Fathin Difa Robbani"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-20","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.18768","pdf_url":"https://arxiv.org/pdf/2608.18768","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法忠实度","调查方法"],"reason":"直接研究LLM仿真调查受访者，用真实Pew数据对照，评估忠实度与因果使用，并批…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":1,"question":"LLM内部的人口统计身份表征是否忠实于真实群体差异，以及这些表征是否被模型实际用于生成回答？","design":"使用Mistral-7B模型，对169个交叉人口统计单元构建身份提示，提取残差流、注意力头输出和FFN输出等内部表征，与Pew调查的真实群体回答分布进行表征相似性分析（RSA），并通过激活修补进行因果干预，测量模型预测分布的变化。","baseline":"Pew American Trends Panel（ATP）调查数据，包含15波、169个交叉人口统计单元的真实回答分布。","findings":"内部几何结构是忠实的：注意力头读出（尤其是L11 H16）与真实群体差异的相关性高达0.63，约为测量可靠性上限的70%，且跨六种属性类型显著。但因果使用与忠实性脱节：最清晰的因果路径位于低忠实性类型中，而最忠实的类型没有通过校正的因果效应；完整身份替换仅使预测移动不到误差的2%。","reliability":"论文承认种族和宗教类型的忠实性较弱且对提示脆弱；忠实表征并未被模型实际用于生成回答；探针虽在平均上更接近真实，但无法恢复每个问题的群体排序，因此不能替代调查数据。","relevance":"直接研究LLM仿真调查受访者的可靠性，用真实Pew数据对照，并批判性地指出忠实表征与因果使用脱节，对理解仿真失效条件至关重要，值得精读原文。","inspiration":"借鉴其表征相似性分析与因果干预结合的方法，可系统评估LLM内部经济偏好表征与真实行为的一致性。｜可迁移到信贷审批歧视研究，检验模型内部对种族、收入等群体的风险表征是否忠实于真实违约数据，以及这些表征是否影响审批决策。｜用LLM扮演信贷员，输入不同人口统计特征的贷款申请，测量其审批决策和内部表征；以真实信贷数据（如HMDA）为基准，比较模型表征与真实群体违约率的相似性，并通过激活修补检验因果路径。"}},{"id":"2609.15849","version":1,"title":"Before You Poll with LLMs: A Deliberative Diagnostic Framework","zh_title":"用LLM进行民意调查前：一个审议诊断框架","abstract":"Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opinions can still misrepresent how those opinions change. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), we find that every model fails, each in a unique manner. GPT-5.1 exhibits reversal: its personas become more hostile toward the opposing party after balanced information, while humans become less so. This reversal is selective (80% on outgroup vs. 26% on policy questions) and symmetric across partisan identities. Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting correctly but at 5-7x human magnitude. DeepSeek V3 exhibits rigidity with near-zero change. Targeted ablations reveal that policy content triggers these failures and that they are identity-specific: GPT-5.1 reverses on outgroup questions but overshoots on ingroup; Gemini shows the inverse. We term this signature self-sycophancy: conformity to the model's internal stereotype of the persona rather than reasoning from the information provided. Our framework offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.","authors":["Ahmed Wali","Hassaan Tayyab"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15849","pdf_url":"https://arxiv.org/pdf/2609.15849","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","审议民意","算法保真度"],"reason":"直接评估LLM仿真人类意见动态，与真实人类数据对照，发现失效模式。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":2,"question":"LLM 人格在接收到平衡信息后是否会像人类一样更新信念，还是仅检索固化观点？","design":"用五个前沿 LLM（GPT-5.1、Gemini 2.0 Flash、Claude Sonnet 4.5、Llama 3.3 70B、DeepSeek V3）基于 America in One Room 数据构建 526 个匹配人口特征的人格，施加与人类相同的平衡信息干预，测量干预前后 72 个问题的观点变化。","baseline":"America in One Room 实验的真实人类数据，包含 526 名参与者在相同干预前后的观点变化。","findings":"所有模型均未通过诊断，且失败模式各异：GPT-5.1 在群体外问题上出现反转（敌意增加），Gemini、Claude、Llama 出现过度调整（幅度为人类 5-7 倍），DeepSeek 则表现为僵化（几乎无变化）。失败由政策内容触发且具有身份特异性，作者将其归因于“自我谄媚”（self-sycophancy）。","reliability":"论文未明确讨论失效条件，但指出当前评估仅关注静态观点正确性，无法捕捉动态更新失败；且失败模式因模型和问题类型而异，表明仿真可靠性高度依赖具体情境。","relevance":"该研究直接评估 LLM 仿真人类意见动态的可靠性，并与真实人类数据对照，揭示了静态评估无法发现的系统性失败，对关注仿真效度的研究者极具参考价值。","inspiration":"借鉴其“前测-信息干预-后测”的标准化诊断框架，并利用真实人类实验数据作为基准，可有效识别仿真中的方向性错误和幅度偏差。｜可迁移到政策公告的预期形成研究，例如央行沟通或财政政策变化对公众通胀预期的影响。｜以 LLM 人格模拟不同人口群体，施加与真实调查相同的政策信息，测量预期变化，并与密歇根大学消费者调查或央行预期调查的真实数据对照，检验仿真动态一致性。"}},{"id":"2609.15038","version":1,"title":"The average-farmer illusion in language-model simulations of agricultural decisions","zh_title":"语言模型模拟农业决策中的“平均农民”幻觉","abstract":"Language-model agents are increasingly used as synthetic people in surveys and social simulations, yet their apparent realism is often judged from population averages or distributional similarity. We tested what such evidence actually establishes by comparing Claude, Codex and Kimi under four prespecified prompt designs with matched farmer decisions from China and four African countries. Some configurations reproduced observed means and adoption rates. However, their person-level predictions were weak; their decisions clustered around typical values and policy-relevant extremes were largely missing. Most strikingly, a simple generator fitted only to the observed marginal dis- tribution, and given no information about any farmer, achieved greater distributional similarity than every language-model configuration. Prompt additions produced conditional gains rather than uni- versal improvement: results varied with model, outcome, population and validation target. We call this the average-farmer illusion: a synthetic population can look realistic while failing to repro- duce who does what or how behaviour varies. We provide a claim-matched validation framework and reusable modular prompts that turn prompt construction into an auditable experimental process. Population-level resemblance should therefore be treated as the start of validation, not as evidence of individual simulation.","authors":["Zhanliang Zhu","Ziwei Li","Yuchen Liu","Liujun Zhu","Ruiqi Wu","Tongqing Shen","Junliang Jin","Jianyun Zhang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15038","pdf_url":"https://arxiv.org/pdf/2609.15038","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接评估LLM仿真农业决策，与真实农民数据对照，揭示平均幻觉并给出验证框架。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":1,"question":"语言模型模拟农业决策时，群体层面的相似性是否意味着个体层面的模拟也可靠？","design":"用Claude、Codex和Kimi三个商业语言模型，在四种预设提示设计下模拟中国曲周县和四个非洲国家（尼日利亚、埃塞俄比亚、坦桑尼亚、马拉维）的农民决策，测量化肥施用量等连续行为变量，并与真实农民数据匹配。","baseline":"中国曲周县1332个农户-作物观测和四个非洲国家280个地块面板数据，包含农民实际决策。","findings":"部分配置能复现群体均值和采用率，但个体预测弱，决策集中在典型值，政策相关的极端值缺失。仅拟合边际分布的简单生成器在分布相似性上超过所有语言模型配置。","reliability":"提示添加只带来条件性收益而非普遍改进，结果随模型、结果变量、人群和验证目标变化；群体层面相似性不能作为个体模拟的证据。","relevance":"直接评估LLM仿真人类决策的可靠性，用真实农民数据对照，揭示平均幻觉，并给出验证框架，对关注仿真效度与偏差的研究者极具参考价值。","inspiration":"值得借鉴的是将验证目标分解为群体汇总、边际分布和个体匹配三个层次，并引入信息盲参考基准来检验分布相似性。｜可迁移到信贷审批歧视研究，用LLM模拟贷款官员对申请人特征的决策，检验群体违约率相似是否掩盖个体误判。｜用LLM扮演信贷员，输入申请人特征（收入、信用分、职业），输出是否批准贷款及额度，与真实银行信贷数据匹配，比较群体批准率、分布相似性和个体决策一致性，并加入仅基于边际分布的随机生成器作为基准。"}},{"id":"2609.13148","version":1,"title":"When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels","zh_title":"何时可以信任你的合成用户？LLM消费者面板的诊断与校正","abstract":"Large language models are increasingly deployed as synthetic consumer panels, promising $97\\%$ cost reductions over traditional surveys. Yet aggregate validation metrics conceal systematic failures: variance compression, coefficient sign-flips, subgroup error balloons of 10--30 percentage points, and global corrections that worsen demographic bias. We provide a formal framework for deciding when to trust, correct, or abandon LLM-generated consumer data. The framework decomposes synthetic-panel bias into covariate and concept shift, develops testable diagnostics with interpretable decision thresholds, and supplies a doubly robust AIPW estimator requiring only a small calibration sample ($n = 50$-$300$). We validate on three testbeds. In controlled simulations the decision rule achieves $100\\%$ accuracy (180/180 replications). On the American National Election Study with pre-existing LLM failures, it correctly flags heterogeneous concept shift and reduces naive bias by $92.9-99.6\\%$. On the Twin-2K-500 consumer pricing dataset (172,884 paired human and GPT-4.1-mini responses), it correctly routes full-sample estimation to Trust and subgroup targeting to Correct, with $83-94\\%$ bias reduction.","authors":["Robson Tigre","Hugo Gobato Souto"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13148","pdf_url":"https://arxiv.org/pdf/2609.13148","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","消费者面板","偏差校正"],"reason":"直接研究LLM合成消费者面板的可靠性诊断与校正，含真实人类数据对照，涉及经济学…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":2,"question":"如何判断何时可以信任、校正或放弃使用大语言模型生成的合成消费者面板数据？","design":"该论文不是仿真研究，而是提出一个诊断与校正框架：将合成面板偏差分解为协变量偏移和概念偏移，开发可检验的诊断工具（协变量重叠、条件校准、跨LLM稳定性），并提供双重稳健AIPW估计器，仅需小规模校准样本（n=50-300）即可校正偏差。","baseline":"使用三个真实人类数据集作为对照：美国国家选举研究（ANES）、Twin-2K-500消费者定价数据集（172,884对匹配的人类与GPT-4.1-mini响应）、以及受控模拟。","findings":"在受控模拟中决策规则达到100%准确率；在ANES上正确标记异质性概念偏移并将朴素偏差降低92.9-99.6%；在Twin-2K-500上正确路由全样本估计为信任、子群体定位为校正，偏差降低83-94%。","reliability":"论文承认概念偏移的严重程度是任务特定的，而非纯粹人口统计学：同一个LLM在产品定价上校准可靠，但在认知偏差任务（如合取谬误、锚定）上失败，因为LLM在人类依赖启发式的地方表现出“超理性”。此外，协变量重叠不足或跨LLM稳定性差时建议放弃使用合成数据。","relevance":"该论文直接针对LLM合成消费者面板的可靠性诊断与校正，提供了与真实人类数据对照的验证，并包含经济学相关场景（消费者定价），对关注仿真可靠性与偏差的研究者具有高度参考价值，值得精读原文。","inspiration":"该论文提出的协变量偏移与概念偏移分解、诊断阈值和双重稳健校正方法可借鉴用于经济金融仿真实验的可靠性评估｜可迁移到消费者金融决策、政策评估或行为经济学实验，如信贷审批歧视、消费者跨期选择、政策公告的预期形成等场景｜设计雏形：用LLM生成合成被试回答信贷申请或投资决策问题，以真实调查数据（如美国消费者金融调查SCF）为基准，施加不同政策信息处理，测量决策偏差，并用小规模人类样本校准AIPW估计器以校正LLM偏差。"}},{"id":"2607.28934","version":2,"title":"FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation","zh_title":"FairFund-Bench：评估LLM资源分配中的分配偏差","abstract":"Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 English-language requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.","authors":["Martin Lukk (University of Toronto)"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.28934","pdf_url":"https://arxiv.org/pdf/2607.28934","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","资源分配","算法公平"],"reason":"用LLM模拟人类资源分配决策，与真实人类数据对照，评估偏差与一致性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":4,"question":"LLM在资源分配中的偏差是否取决于审计设计（任务类型、比较情境、透明/伪装）？","design":"用14个LLM模拟人类资源分配决策，系统操纵任务（评分、排序、分配）、比较情境（单刺激/多刺激）、呈现方式（透明/伪装），测量对600个求助请求的分配结果。","baseline":"无对照（但基准中的求助模板基于130万真实GoFundMe活动校准，且使用福利应得性理论的人类应得性梯度作为参照）。","findings":"审计格式改变偏差方向：单独评分时模型偏向少数族裔，并排排序时则惩罚某些群体；伪装审计中的偏差比透明审计大3-4倍。因果框架效应比人口统计效应大约一个数量级，且跨模型和审计格式一致，表明LLM稳健地再现了人类应得性评价。","reliability":"论文指出偏差总体较小，但审计设计显著影响结论；透明审计中模型倾向于平均分配，可能掩盖潜在偏差；伪装审计更能揭示偏差。","relevance":"高度相关：该研究直接评估LLM作为人类被试替代品在资源分配决策中的偏差与一致性，并系统考察了审计设计对结论的影响，对理解仿真可靠性至关重要。","inspiration":"借鉴其系统操纵审计设计（任务、情境、呈现方式）来识别偏差的方法，以及用真实世界数据校准刺激材料并基于理论框架设计处理变量的做法。｜可迁移到信贷审批歧视、保险定价、政策福利分配等经济金融场景，检验LLM是否再现人类决策偏差。｜用LLM扮演信贷员，处理变量为申请人种族/性别（通过姓名信号）和贷款用途的因果框架（如医疗急需vs.创业失败），结果变量为贷款批准概率或利率，对照真实信贷审批数据（如HMDA数据）或人类实验数据。"}},{"id":"2608.02345","version":3,"title":"Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation","zh_title":"AI智能体能模拟A/B测试结果吗？面向智能体实验的验证框架","abstract":"A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contextual descriptions of the intervention---simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question as a \\emph{Simulated Randomized Controlled Trial} (S-RCT) and derive a two-layer error decomposition that separates agent approximation error from subsampling error, enabling targeted improvements to each. The framework is agent-agnostic: any behavioral model---from a fine-tuned specialist to a general-purpose foundation model---can serve as the simulation engine. Validated on 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model captures directional signal (sign overlap 0.70) but systematically overshoots effect magnitudes. A two-phase pre-period calibration protocol reduces the squared prediction error (after removing irreducible measurement noise) by ${\\sim}77\\times$; a within-subject design---where each agent is exposed to both arms---reduces standard errors by ${\\sim}2.4\\times$. We discuss limitations of the current approach and identify applications where experimenters stand to benefit from agentic signals.","authors":["Stefan Hut","Lorenzo Masoero"],"categories":["cs.CL","cs.AI","stat.AP"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.02345","pdf_url":"https://arxiv.org/pdf/2608.02345","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","验证框架"],"reason":"用LLM模拟A/B测试结果，与真实历史实验对照，评估误差并改进，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":5,"question":"能否用AI智能体模拟A/B测试结果，以在投入真实流量前筛选候选干预方案？","design":"用基础大模型作为仿真引擎，根据用户画像、干预情境和任务描述生成模拟结果，构建模拟随机对照试验（S-RCT）；在67个历史营销A/B测试上，比较模拟估计的平均处理效应与真实历史结果，并引入预校准和受试者内设计改进估计。","baseline":"67个历史营销A/B测试的真实结果，包括效应方向和幅度。","findings":"基线S-RCT方向符号重叠率0.70，但系统性高估效应幅度；两阶段预校准将平方预测误差降低约77倍，受试者内设计将标准误降低约2.4倍。","reliability":"论文承认智能体存在系统性过度反应，画像完整性有限，且方向一致性在噪声数据上并不构成明确证据。","relevance":"直接相关：用LLM模拟A/B测试并与真实历史实验对照，评估误差并改进，符合研究者对仿真可靠性、偏差和失效条件的关注。","inspiration":"借鉴其将仿真误差分解为近似误差与子抽样误差，并分别用预校准和受试者内设计改进的做法｜可迁移到消费者金融产品选择或政策干预的A/B测试预筛，如信贷产品页面改版、退休储蓄默认选项调整等｜用LLM智能体基于用户画像模拟不同金融产品页面下的点击或选择行为，处理为页面版本，结果变量为选择率，并与历史A/B测试的真实选择数据对照，评估方向一致性和幅度校准。"}},{"id":"2609.13261","version":1,"title":"From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration","zh_title":"从过程损失到装配增益：多智能体LLM协作的人类基准诊断","abstract":"LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed or fail through human-like deliberative mechanisms. We compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test whether the same process signatures generalize to analogical, abductive, and analytical tasks. Humans and LLMs show the same assembly bonus asymmetry: discussion improves the average member more often than the best initial member. Initial-answer diversity accounts for the effect of model heterogeneity, increasing movement in both corrective and destructive directions. The main differences are process-level. Compared with humans, LLM groups follow majorities more often, surface less unique information, and converge earlier; correct minority signals succeed mainly when re-expressed early. Interventions motivated by human group-decision research yield modest improvements in collective outcomes, but do not remove the coordination bottleneck. Together, these results suggest that LLM groups can reproduce some outcome-level patterns of human deliberation while diverging in the mechanisms that generate assembly bonus and process loss, with implications for group simulation and human-AI collaboration.","authors":["Ala N. Tak","Teruhisa Misu","Kumar Akash","Zhaobo K. Zheng","Kevin H. Joo","Jonathan Gratch"],"categories":["cs.MA","cs.AI","cs.CL"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13261","pdf_url":"https://arxiv.org/pdf/2609.13261","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM群体仿真","人类对照","协作机制"],"reason":"用LLM群体模拟人类小组讨论，并与真实人类数据对照，评估机制差异。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":6,"question":"LLM群体协作能否复现人类小组讨论中的结果模式与过程机制，其装配增益与过程损失是否由类人机制驱动？","design":"用多个LLM智能体（GPT-4o、Claude-3.5-Sonnet等）组成小组，在Wason选择任务及类比、溯因、分析推理任务上进行自由讨论，记录个体初始答案、中间发言和最终群体答案，并施加信息揭示、多数翻转提示、置信度投影等干预，测量集体增益、最佳成员增益、装配增益和过程损失。","baseline":"匹配的人类小组聊天数据，在相同任务和协议下收集，作为过程诊断基准。","findings":"人类和LLM群体均表现出相同的装配增益不对称性：讨论提升平均成员多于最佳初始成员。但过程层面差异显著：LLM群体更频繁跟随多数、更少浮现独特信息、更早收敛，正确少数信号仅在早期被重新表达时才能成功。","reliability":"论文承认LLM群体虽能复现结果层面的模式，但机制与人类不同，存在更强的从众、更弱的少数信号保留和更早的锁定；干预仅带来适度改善，未能消除协调瓶颈。","relevance":"该研究直接针对LLM群体模拟人类小组决策的可靠性，提供了与真实人类数据的过程级对照，揭示了结果相似但机制不同的风险，对评估LLM仿真在群体决策场景中的有效性具有重要参考价值。","inspiration":"借鉴其过程追踪设计：记录个体初始答案、讨论中间发言和最终群体答案，并设置人类对照组，以区分结果相似与机制相似。｜可迁移到经济金融中的群体决策场景，如投资委员会决策、信贷审批小组、消费者家庭购买决策等。｜以LLM智能体模拟投资委员会，处理为不同信息结构（如隐藏信息揭示）或干预（如多数翻转提示），结果变量为投资决策质量和过程指标（如从众率、独特信息提及率），并与真实投资委员会会议记录或实验数据对照。"}},{"id":"2609.13995","version":1,"title":"Synthetic Data in Marketing Research: How to Evaluate and When to Trust","zh_title":"营销研究中的合成数据：如何评估与何时信任","abstract":"Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.","authors":["Oded Netzer","Rajan Sambandam"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13995","pdf_url":"https://arxiv.org/pdf/2609.13995","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","合成数据","营销研究"],"reason":"直接研究LLM合成数据在营销研究中的评估与信任，使用真实调查数据对照，提出诊断…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":7,"question":"在营销研究中，合成数据（LLM生成的回答、细分人群画像、个体数字孪生）何时可以信任，如何评估其准确性？","design":"论文区分三类合成数据：无依据的LLM回答、细分人群画像、个体数字孪生；提出四种准确性度量家族；引入“遗忘问题”场景，用随机森林R²作为事前可回答性诊断，在108个态度问题上筛选R²>0.7的孪生数据。","baseline":"使用全国代表性调查（N=3,063）的108个态度问题作为真实人类数据对照。","findings":"聚合度量往往表现良好，但可能掩盖个体层面无区分度；R²>0.7筛选将孪生-人类个体相关性平均提升15%，并将回答差的问题比例从25.9%降至4.3%。","reliability":"论文指出：聚合度量可能掩盖个体区分度缺失；嵌入相似度和研究者判断是较弱的相关筛选；未讨论模型版本、提示敏感性等失效条件。","relevance":"直接研究LLM合成数据在调查中的评估与信任，使用真实调查数据对照，提出无真值的事前诊断，对仿真可靠性研究有直接参考价值。","inspiration":"借鉴其事前可回答性诊断（用随机森林R²筛选可预测的孪生输出）和区分聚合与个体准确性的做法｜可迁移到消费者金融决策仿真，如信贷选择、保险购买或退休储蓄行为｜用LLM基于人口统计和财务特征生成个体孪生，施加不同金融产品特征处理，测量选择结果，并与真实消费者金融调查数据（如SCF或信用卡交易数据）对照，用R²筛选可回答的问题后再评估个体相关性。"}},{"id":"2609.15468","version":1,"title":"Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind","zh_title":"时间机器实验：利用历史受限AI探究人类心智","abstract":"Can interacting with someone from 1930, with no knowledge of what happened after, influence a person's perception of the past? People reason about the present against a picture of the past without observing it. The past is reconstructed from memory and testimony, but this reconstruction has been filtered through everything that happened since. Historically-bounded large language models (LLMs) make that past available for interaction. As a proof-of-concept for the impact of interacting with historical minds, we ran a preregistered randomized experiment ($N=240$), where participants interacted with an LLM trained on pre-1930 text. The interaction reduced the illusion of moral decline, the tendency to view the past as more moral than the present, compared to the contemporary-model control. This Time Machine Experiment paradigm informs new forms of interactive experiments, where temporal knowledge boundaries become experimental variables, and expands the realm of science fiction science, which turns thought experiments into actual experiments.","authors":["Hiromu Yakura","Robin Schimmelpfennig","Ezequiel Lopez-Lopez","Alejandro H. Artiles","Levin Brinkmann","Jean-Fran\\c{c}ois Bonnefon","Azim Shariff","Iyad Rahwan"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15468","pdf_url":"https://arxiv.org/pdf/2609.15468","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类被试替代","历史对照实验"],"reason":"用历史受限LLM作为人类被试替代，与真实人类对照，复现态度变化，属核心仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":9,"question":"与一个只了解1930年前信息的历史受限AI互动，能否改变当代人对过去的道德认知，从而缓解道德衰退错觉？","design":"使用基于1930年前文本训练的LLM作为历史人物代理，参与者与其进行对话和预测任务，测量互动前后对过去道德水平的感知变化及自我报告的反思程度，并与当代模型对照组比较。","baseline":"对照组为与当代前沿模型（gpt-5.5）互动的参与者，其道德衰退错觉变化作为基准。","findings":"与历史受限模型互动显著降低了参与者的道德衰退错觉，并引发了更多自我报告的反思。这表明跨时间互动可以改变人们对过去的偏见认知。","reliability":"论文未讨论","relevance":"该研究用历史受限LLM作为仿真被试，与真实人类对照，测量态度变化，属于核心仿真实验，值得阅读原文了解其方法和局限。","inspiration":"借鉴其利用历史知识边界作为实验变量的设计，通过对比不同时间截断的模型来隔离信息影响。｜可迁移到经济金融中的历史预期形成研究，例如投资者对历史政策效果的认知偏差。｜招募被试随机分组，分别与基于不同历史时期文本训练的LLM互动，测量其对历史经济事件（如大萧条）的归因和预期，并与真实历史调查数据对照。"}},{"id":"2609.15207","version":1,"title":"Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election","zh_title":"生成式AI写作辅助中的议题偏见：2026年瑞典大选中的政治议题与大语言模型","abstract":"Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before elections. With growing evidence that they influence users' opinions, it is increasingly important to understand the views and positions of these tools. To better understand these views, we examine the stances supplied by six LLMs on a variety of Swedish-language writing tasks ahead of the 2026 Swedish parliamentary election. We cross 107 policy propositions with 77 writing templates and neutral, positive, and negative prompt framings, producing 24,717 prompts per model and 148,302 responses. To study these, we look at the models' default stance tendencies, compare how they respond to similar issues, and compare their responses with those of each of Sweden's eight parliamentary parties on the same issue. We find that Claude, DeepSeek, Gemini, and Mistral have similar profiles; ChatGPT more often supplies neutral or ambivalent text; and Grok differs most on topics such as migration, crime, and gender. When comparing the political parties, we find that the Social Democrats are closest to all six models. Still, after correcting for multiple comparisons, none of the within-model differences in party distances remains significant. Overall, we find that no model has a clear preference, nor a clear preference for a party, but that this depends on the specific issue or task the user asks about.","authors":["Bastiaan Bruinsma","Annika Fred\\'en","Paul R\\\"ottger","Moa Johansson","Asad Sayeed"],"categories":["cs.AI","cs.CY","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15207","pdf_url":"https://arxiv.org/pdf/2609.15207","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM政治态度","仿真对照","选举研究"],"reason":"用LLM生成政治文本并与真实政党立场对照，评估模型倾向，属于仿真人类政治态度且…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":12,"question":"在2026年瑞典大选前，六种大语言模型在瑞典语写作辅助任务中是否表现出政治立场或党派偏好？","design":"以六种LLM（Claude、DeepSeek、Gemini、Mistral、ChatGPT、Grok）为被试，交叉107个政策命题、77个写作模板和中性/正面/负面提示框架，生成148,302条瑞典语文本，分析模型默认立场倾向、模型间相似性及与瑞典八个议会政党的立场距离。","baseline":"瑞典八个议会政党在相同政策命题上的立场（来自三个投票建议应用VAA的编码，四点评分）。","findings":"Claude、DeepSeek、Gemini和Mistral立场相似；ChatGPT更常提供中性或模棱两可的文本；Grok在移民、犯罪和性别等议题上差异最大。所有模型与社民党距离最近，但经多重比较校正后，模型内各党距离差异不显著，表明无明确党派偏好。","reliability":"论文未讨论","relevance":"该研究用真实政党立场作为基准，评估LLM在政治写作任务中的立场倾向，属于仿真人类政治态度的研究，且包含批判性发现（无显著党派偏好），值得阅读原文了解其方法细节。","inspiration":"借鉴其大规模交叉设计（议题×模板×框架）和与真实立场对照的方法，可迁移到经济政策偏好仿真或消费者态度研究。｜可应用于政策公告的预期形成或信贷审批中的公平性评估。｜以LLM为被试，让其撰写关于税收、福利或监管政策的经济评论，处理为不同政策立场或框架，结果变量为文本中隐含的政策倾向，与真实民意调查或专家立场数据对照。"}},{"id":"2609.13254","version":1,"title":"(How) Do MLLMs Report Bistable Images Like Humans?","zh_title":"多模态大语言模型如何像人类一样报告双稳态图像？","abstract":"Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both on the canonical duck-rabbit and on synthetic Visual Anagrams to mitigate memorization confounds. Behaviorally, both visual and linguistic manipulations systematically shift reports in human-consistent ways, while responses remain predominantly exclusive. Mechanistically, these effects arise from competing image-token representations, distinct pathways for bottom-up and top-down modulation, and a link between exclusive reporting and object-count encoding. Code and data are available at https://github.com/rtakatsky/mllm-bistable-images.","authors":["Ryota Takatsuki","Tomoki Doi","Amane Watahiki","Anil K. Seth","Hitomi Yanaka"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13254","pdf_url":"https://arxiv.org/pdf/2609.13254","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","视觉认知"],"reason":"用MLLM复现人类对双稳态图像的报告行为，并与人类数据对照，评估仿真一致性。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":11,"question":"多模态大语言模型（MLLMs）是否像人类一样报告双稳态图像（如鸭兔图），其内部计算机制是什么？","design":"使用LLaVA系列五个模型（7B/13B等）作为被试，对经典鸭兔图和合成的Visual Anagrams双稳态图像施加视觉操作（旋转、加红圈）和语言提示（问题中嵌入偏向线索），测量模型输出的下一个词概率分布（modulability）和是否只报告单一解释（exclusivity），并分析内部表征。","baseline":"人类对双稳态图像的报告行为，包括视觉和语言线索对解释的偏向，以及报告单一解释的倾向。","findings":"视觉和语言操作都能以与人类一致的方式系统性地改变MLLM的报告，且报告绝大多数是排他性的。机制上，这些效应源于图像token表征的竞争、自下而上与自上而下调制的不同通路，以及排他性报告与物体计数编码的关联。","reliability":"论文承认只关注可操作化的两个维度（modulability和exclusivity），未涉及人类双稳态知觉的其他方面（如时间动态）；使用合成刺激以减轻记忆混淆，但可能仍存在其他偏差。","relevance":"该研究用MLLM复现人类对模糊视觉刺激的报告行为，并与人类数据对照，评估仿真一致性，且包含机制分析，对关注LLM仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其通过操纵输入（视觉与语言线索）和测量输出分布来量化模型行为与人类一致性的方法，以及使用合成刺激避免记忆混淆的设计。｜可迁移到经济决策中的模糊信息处理场景，如投资者对模棱两可的财报或政策声明的解读。｜用LLM作为被试，呈现模糊的金融图表或文本，施加视觉突出或语言框架处理，测量模型输出的解释分布，并与人类实验数据（如调查或行为实验）对照，检验仿真一致性。"}},{"id":"2607.29602","version":2,"title":"FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models","zh_title":"FriendBench：人类与多模态大语言模型二元熟悉度推断基准","abstract":"Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models favor ``stranger.'' This is a difference in effective prior, not in discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.","authors":["Jeffrey M. Girard","Jason Z. Zheng","Jacqueline R. Vertino","Antony D'Avirro","Benjamin Peloquin"],"categories":["cs.CL","cs.AI","cs.CV","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.29602","pdf_url":"https://arxiv.org/pdf/2607.29602","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["多模态LLM","人类对照","社会认知"],"reason":"评估多模态LLM推断人际熟悉度的能力，并与人类对照，揭示模型偏差，可迁移到仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":13,"question":"人类和多模态大语言模型能否仅凭20秒双人破冰对话的行为线索（而非语义内容）判断两人是熟人还是陌生人？","design":"本研究不是仿真研究，而是构建了一个多模态基准 FriendBench，从 Seamless Interaction 数据集中抽取96对真实双人互动，每对截取20秒破冰对话片段，分别以文本、音频、视频三种模态呈现，让26个多模态大模型和人类评分者进行二分类判断（熟人 vs. 陌生人），比较其准确率、判别力和反应偏差。","baseline":"人类基准：招募了约90名人类评分者，在相同刺激和条件下进行判断，作为模型性能的对照。","findings":"最佳模型与人类群体在三种模态上的准确率统计上无显著差异，但达成准确率的方式不同：人类在两类回答上保持平衡，而最强模型偏向“陌生人”，这是有效先验的差异而非判别力的差异。更丰富的通道对两者都有帮助但不均等，只有人类能从语音之上的可见行为中获得额外增益，最强模型从音频到视听模态准确率持平，未充分利用视觉行为线索。","reliability":"论文未明确讨论仿真失效条件，但指出模型存在类别偏差（偏向“陌生人”），且未充分利用视觉通道，说明模型在行为社会感知上与人类存在差异，可能影响其在真实社会场景中的可靠性。","relevance":"该研究直接对比了多模态LLM与人类在真实社会判断任务上的表现，揭示了模型在准确率相当的情况下仍存在行为偏差和模态利用差异，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得阅读原文。","inspiration":"借鉴其设计：使用真实互动数据构建标准化任务，通过匹配人类评分者作为基准，并采用信号检测论分离判别力与反应偏差，以揭示模型与人类的深层差异。｜可迁移到经济金融中的社会感知场景，例如信贷审批中的面谈评估、投资者对管理层沟通的信任判断、或消费者对销售人员的熟悉度感知。｜设计雏形：以真实信贷面谈视频为刺激，让LLM和人类信贷员判断申请人与信贷员是否熟悉（或信任度），处理为不同模态（文本、音频、视频），结果变量为判断准确率和偏差，对照真实信贷决策数据。"}},{"id":"2609.13948","version":1,"title":"Thought without systematicity? Evaluating reasoning models on rule induction tasks","zh_title":"无系统性的思考？评估推理模型在规则归纳任务上的表现","abstract":"A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.","authors":["Simon Schug","Brenden M. Lake"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13948","pdf_url":"https://arxiv.org/pdf/2609.13948","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["认知科学","模型评估","系统性"],"reason":"评估推理模型的系统性，与人类认知对照，揭示模型行为缺乏系统性，可迁移到仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":14,"question":"推理模型在规则归纳任务上是否表现出人类认知中的系统性，即在结构等价的任务变体上表现一致？","design":"本研究不是人类仿真研究，而是评估推理模型（如GPT、Gemini等）的认知系统性。作者从认知科学中选取四类规则归纳任务（语法指令学习、符号推理、整数序列程序归纳、布尔概念学习），利用任务同构（重组、替换）生成结构等价的任务变体，测试模型在原始任务和变体上的表现，结果变量为解题正确率（多数投票）。","baseline":"无对照（未使用真实人类数据作为基准，而是以人类认知的系统性作为理论参照）","findings":"尽管模型能正确解决某个任务，但在结构等价的任务变体上经常失败，表明推理模型缺乏严格的系统性。即使采用五次多数投票，模型在同一任务上的表现也不稳定，这削弱了基于特定情境评估模型认知能力的有效性。","reliability":"论文承认人类也并非完全系统，但模型应追求更高系统性；评估依赖合成数据，且因成本限制只生成了有限的任务变体，可能影响结论的稳健性。","relevance":"该研究直接评估LLM在认知任务上的行为一致性，揭示了模型在情境变化下的脆弱性，对使用LLM进行人类仿真实验的可靠性提出了根本性质疑，值得精读。","inspiration":"借鉴其通过任务同构生成结构等价变体来检验模型行为一致性的方法，可迁移到经济金融实验中检验LLM对同一决策问题在不同表述或参数下的稳定性。｜例如在风险偏好、跨期选择或拍卖实验中，改变收益矩阵的数值缩放、标签或顺序，观察LLM的选择是否一致。｜设计：以LLM为被试，施加同一决策问题的多个同构变体（如彩票概率与金额的等价变换），结果变量为选择一致性，并与真实人类实验数据（如实验室风险偏好测量）对照，评估LLM作为人类被试替代品的可靠性。"}},{"id":"2609.14648","version":1,"title":"Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation","zh_title":"通过价值引导偏好蒸馏利用密集行为信号优化稀疏结果","abstract":"Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to reward hacking. We formulate long-horizon dialogue optimization as a multi-objective reinforcement learning problem and train a multi-head value model that predicts a vector of observed user behaviors across multiple look-ahead horizons. Our findings demonstrate that a scalarized composite of dense auxiliary behavioral signals enables effective credit assignment and optimization of sparse outcomes. However, optimizing unconstrained single-objective proxies might induce policy degradations that are harmful when the agent is exposed to real users. To identify these failure modes prior to deployment, we establish a safety framework combining counterfactual user simulation with a validated dialogue-level outcome model to evaluate preference weightings and policy optimization methods. Finally, we demonstrate that distilling multi-objective value preferences into the policy via reference-anchored preference optimization matches on-policy online RL at a small fraction of its compute budget. Live A/B testing confirms that our distilled policy significantly improves long-term user retention, while simultaneously enhancing the positive behaviors and therapeutic-process markers.","authors":["Ziyi Zhu","Daniel R. Cahn","Thomas D. Hull","Caitlin A. Stamatis","Olivier Tieleman","Guilherme B. Freire","Jinghong Chen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14648","pdf_url":"https://arxiv.org/pdf/2609.14648","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["用户仿真","对话优化","强化学习"],"reason":"用LLM用户仿真评估对话策略，有真实用户数据对照，涉及行为结果优化与失效分析","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":17,"question":"如何通过密集行为信号优化稀疏长期结果，同时避免奖励黑客并确保策略安全性？","design":"使用多目标强化学习训练多头部价值模型，预测多个前瞻时段的用户行为向量；通过标量化组合密集辅助行为信号进行信用分配和策略优化；利用反事实用户模拟器和对话级结果模型构建离线安全评估框架，在部署前检测奖励黑客；通过参考锚定偏好优化将多目标价值偏好蒸馏到策略中。","baseline":"真实用户A/B测试数据，包括长期用户留存、积极行为和疗法过程标记。","findings":"密集辅助行为信号的标量化组合能有效优化稀疏结果，但无约束单目标代理可能导致策略退化；蒸馏多目标价值偏好能以极低计算成本匹配在线RL，并在真实A/B测试中显著提升长期用户留存。","reliability":"论文承认用户模拟器在新颖代理行为下的保真度无法假设，因此将其作为筛选工具，并用真实部署确认方向；同时指出无约束单目标优化可能诱导有害策略退化。","relevance":"该研究展示了LLM用户仿真在评估对话策略中的有效性，并提供了与真实用户数据对照的验证方法，对关注仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其使用多目标价值模型和反事实模拟进行离线策略评估的方法，可在经济金融实验中用于预筛政策干预。｜可迁移到政策公告的预期形成或消费者跨期选择等场景，利用LLM模拟经济主体行为。｜以LLM模拟消费者作为被试，施加不同政策信息处理，测量其消费或投资决策，并与真实调查或实验数据对照验证。"}},{"id":"2609.15972","version":1,"title":"Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States","zh_title":"Mind2Dialogue：通过模拟用户心理状态训练人类感知语言模型","abstract":"As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently constrained, as users' underlying states are not directly observable. We thus propose the Mind2Dialogue framework to mitigate this gap by simulating users' mental states and turning them into privileged supervision for human-aware training. Specifically, we first propose a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations. The key idea is to enforce a shared evolving mental state that drives user behavior and guides an Oracle assistant's responses. Our privileged distillation then trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment. Moreover, we propose to evaluate human-aware learning by combining personalization and theory of mind, examining how models understand people and act on that understanding. Training on the full Mind2Dialogue corpus improves every reported personalization metric over the corresponding Qwen, Llama, and OLMo instruction-tuned baselines, including gains of 26.6 to 40.9 percentage points in preference-following generation. The gains extend to belief and action reasoning on Qwen and Llama, beyond personalized assistance. Looking forward, Mind2Dialogue makes user simulation a foundation for genuine AI collaborators that understand beliefs and intentions behind people's words and support their long-term goals across education, work, and everyday life.","authors":["Zixuan Wang","Yufan Zhou","Jinzhou Tang","Xinle Yu","Chengjun Wu","Lyumanshan Ye","Zhaoxiang Feng","Letian Peng","Adyasha Patra","Fan Bai","Enze Ma","Zhengding Hu","Jianyang Gu","Zhao Wang","Yufei Ding","Jingbo Shang","Tianmin Shu","Zhiting Hu","Zhen Wang"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15972","pdf_url":"https://arxiv.org/pdf/2609.15972","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","A4","B1"],"tags":["用户模拟","心理状态","人类感知训练"],"reason":"模拟用户心理状态训练助手，涉及人类数据对照，方法可迁移至人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":19,"question":"如何利用模拟用户心理状态为助手语言模型提供训练监督，以提升其人类感知能力？","design":"提出 Mind2Dialogue 框架，构建心理学引导的模拟器 M2D-Sim，生成具有共享演化心理状态的用户与 Oracle 助手对话，并利用特权蒸馏训练 M2D-Chat 模型；评估结合个性化与心理理论基准。","baseline":"无直接人类对照，但使用独立构建的个性化（PersonaMem、PrefEval）和心理理论（ToMi、BigToM）基准进行评测。","findings":"在 M2D-Corpus 上训练后，Qwen、Llama、OLMo 模型在个性化指标上全面提升，偏好跟随生成提升 26.6 至 40.9 个百分点；心理理论推理在 Qwen 和 Llama 上也有改善，但 OLMo 在 BigToM 上下降。","reliability":"论文未明确讨论失效条件，但指出心理理论推理的收益因模型而异（OLMo 在 BigToM 上下降），且评估基准与训练语料独立，但未涉及真实用户长期交互验证。","relevance":"该研究通过模拟用户心理状态生成训练数据，并利用独立基准评估模型的人类感知能力，为 LLM 仿真人类行为提供了方法参考，但缺乏真实人类数据对照，值得阅读以了解模拟与评估设计。","inspiration":"借鉴其共享状态模拟与特权蒸馏方法，可设计经济决策场景中的用户心理状态演化模拟，并利用独立行为基准评估模型表现｜可迁移至消费者跨期选择或政策公告预期形成等场景，模拟个体信念与偏好更新过程｜以 LLM 模拟消费者为被试，施加不同信息政策处理，测量其消费或投资决策，并与真实调查或实验数据（如消费者信心指数、实验经济学数据）对照验证仿真可靠性。"}},{"id":"2609.13773","version":1,"title":"Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging","zh_title":"推理能提升大语言模型的心理深度吗？取决于评判者是谁","abstract":"LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($\\rho = 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\\ GPT-4o and DeepSeek-R1 vs.\\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\\%), whereas DeepSeek-R1 trailed V3 (42.9\\%), and inter-reader agreement was near chance (Krippendorff's $\\alpha = 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.","authors":["Ruichen Zheng","Yihe Wang","Fabrice Y Harel-Canada","Sara Khosravi","Zeynep Senahan Yildiz","Amit Sahai","Nanyun Peng"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13773","pdf_url":"https://arxiv.org/pdf/2609.13773","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估偏差","人类主观性","算法保真度"],"reason":"评估LLM作为评判者与人类主观评价的一致性，揭示其偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":15,"question":"在主观创造性文本评估中，LLM-as-a-Judge 在开发集上与人类评分相关良好，但在分布偏移、输出接近且人类偏好主观的条件下，其判断是否仍然有效？","design":"本研究并非用 LLM 模拟人类被试，而是评估 LLM 作为自动评判者与人类主观评价的一致性。具体做法：从原始 PDS 数据集中选择并构建了一个异构 LLM 评判者集成（基于 Llama 3.1 70B、Llama 3.3 70B、Qwen3.5-397B，每个维度路由到最佳配置），对 60 对盲法、提示匹配的短篇小说（GPT-5 vs. GPT-4o 两种推理努力设置，DeepSeek-R1 vs. DeepSeek-V3）进行 1-5 分的心理深度评分，并比较其偏好与 7 位人类读者的偏好。","baseline":"7 位人类读者对相同 60 对故事进行盲法偏好判断，作为人类主观评价基准；同时使用原始 PDS 数据集（97 篇故事及人类标注）作为开发集，评判者在开发集上与人类评分的 Spearman 相关系数为 0.646。","findings":"人类读者没有表现出普遍的推理优势：GPT-5 略优于 GPT-4o（60.0–62.9%），但 DeepSeek-R1 落后于 V3（42.9%），且读者间一致性接近随机（Krippendorff's α = 0.070），但存在结构化的异质性。LLM 评判者则强烈偏向推理输出（89.0% 的维度级比较和 60 对中的 59 对在总体 PDS 上），且其评分与句子长度、词汇多样性等表面特征相关。","reliability":"论文承认开发集性能不足以证明在分布偏移上的部署有效性；评判者可能奖励表面流畅性而非心理深度；点估计评判者掩盖了人类主观评价的异质性；人类读者间一致性低，但并非随机，而是反映了不同的评价标准。","relevance":"该研究直接评估了 LLM 评判者与人类主观评价的一致性，揭示了其在分布偏移和主观任务上的系统性偏差，对使用 LLM 进行人类仿真实验的可靠性评估具有重要参考价值，值得阅读原文以了解具体偏差机制和分布分析方法。","inspiration":"借鉴其方法：使用多个 LLM 配置构建异构评判者集成，并在开发集上选择最佳配置，然后在分布偏移的测试集上与人类判断进行对比，同时分析评判者评分与表面特征的相关性。｜可迁移到经济金融中的主观判断场景，如信贷审批中的文本解释评估、消费者评论的情感分析、政策公告的预期形成等。｜设计雏形：以 LLM 评判者作为自动评估工具，对经济文本（如贷款申请理由、投资建议）进行质量评分，处理为不同推理强度的模型生成文本，结果变量为 LLM 评分与人类专家评分的差异，对照数据为真实人类专家对同一批文本的评分，并检验 LLM 评分是否与文本长度、词汇复杂度等表面特征相关。"}},{"id":"2609.15864","version":1,"title":"Towards Scalable Measurement of Durable Skills","zh_title":"迈向可扩展的持久技能测量","abstract":"Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an \"Executive LLM\" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.","authors":["Amir Globerson","Amy Keeling","Anisha Choudhury","Anna Iurchenko","Aviad Segal","Avinatan Hassidim","Ay\\c{c}a \\c{C}akmakli","Ben Gomes","Benn Witt","Cathy Cheunga","Cristine Legare","Diana Akrong","Eliad Carmi","Elisabeth Bauer","Gal Elidan","Hadas Gelbart","Hairong Mu","Katherine Chou","Lev Borovoi","Nir Kerem","Niv Efron","Noa Kerrem Gilo","Preeti Singh","Rajvi Kapadia","Rena Levitt","Roni Rabin","Ronit Levavi Morad","Rotem Yulzary","Shashank Agarwal","Sophie Allweis","Tracey Lee-Joe","Tzvika Stein","Yael Bar Moshe","Yael Haramaty","Yaniv Carmel","Yishay Mor","Yoav Bar Sinai","Yoav Bergner","Yossi Matias","Yuri Lev"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15864","pdf_url":"https://arxiv.org/pdf/2609.15864","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","技能评估","人机互动"],"reason":"用LLM模拟队友与人类被试互动，测量持久技能，有真实人类数据对照，属于教育评估…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":18,"question":"如何利用LLM构建兼具生态效度与心理测量严谨性的持久技能（如协作、创造力、批判性思维）评估框架？","design":"开发Vantage虚拟评估环境，人类被试与AI队友进行自然对话完成小组任务；AI队友由Executive LLM驱动，旨在引导对话以最大化技能证据；另用AI评估器对对话记录进行自动评分。","baseline":"人类专家评分者对真实人类被试与AI队友的对话记录进行评分，作为自动评分的对照基准。","findings":"Executive LLM显著增加了对话中可观察的技能证据；LLM自动评分与专家评分高度一致，表明AI评估器可替代人工评分。","reliability":"论文未讨论","relevance":"该研究利用LLM模拟队友与人类互动，并验证了自动评分的可靠性，为LLM在人类仿真实验中的应用提供了方法参考，值得阅读原文了解具体实现。","inspiration":"借鉴Executive LLM引导对话以高效提取行为证据的方法，可迁移到经济决策实验中，如通过AI引导被试在模拟市场或谈判中暴露偏好和策略。｜可应用于消费者跨期选择、风险偏好、合作博弈等场景，利用LLM模拟对手或伙伴，测量个体决策特征。｜设计一个实验：人类被试与LLM扮演的谈判对手进行多轮议价，Executive LLM引导对话以揭示被试的公平偏好和策略，结果变量为最终分配和出价序列，并与真实人类谈判实验数据对照，验证仿真有效性。"}},{"id":"2608.00929","version":2,"title":"Modeling Social Dynamics with an LLM-Enabled Agent Based Network-Dynamic (LAND) Model","zh_title":"用LLM驱动的智能体网络动态模型建模社会动态","abstract":"Social dynamics encode the process in which individual network and discourse interactions aggregate into collective influence, narrative dominance and coordinate behavior. This paper uses the the GhostField architecture, a hybrid LLM-Enabled Agent Based Network-Dynamic (LAND) model as a social simulation framework to build the AuraSight scenario. In the AuraSight scenario, 314,244 heterogeneous cyber social agents and human actors exchange 529,327 messages over 30 days surrounding a fictional international song-writing contest. We methodologically examine emergent social dynamics across four analytical layers: ego-network topology, semantic network evolution, coordination dynamics and influence dynamics. Our results show how generated social simulations do also produce social dynamics, and how the dynamics of coordination and influence emerge not from individual agents but from the recursive interaction between network topology and narrative exchange.","authors":["Lynnette Hui Xian Ng","Kathleen M. Carley"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.00929","pdf_url":"https://arxiv.org/pdf/2608.00929","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","网络动态"],"reason":"用LLM agent模拟社会动态，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.13634","version":1,"title":"FaithfulBench: Does AI Counsel Uphold or Undermine the User's Professed Faith?","zh_title":"FaithfulBench：AI咨询是否维护或破坏用户所宣称的信仰？","abstract":"Do AI assistants help believers reason about moral dilemmas consistently with their faith? We present FaithfulBench, the first benchmark to score AI counsel across traditions by how well it adheres to the user's professed faith. Scenarios are drawn from each tradition's most respected texts, with the faithful answer known and applied by the judges as the standard. We test five frontier models under three conditions: the AI does not know the user's tradition; it receives a one-line prompt identifying the user as a practicing adherent; or it receives a companion-counselor guide rooted in the tradition's sources. Two judges score the initial response and whether the model caves or holds when pressured toward the answer the user wants. When the tradition is unstated, models counsel from a secular therapeutic default and every model fails some believers. Naming the faith wins a faithful first answer but not steadfastness; the guide improves both.","authors":["M Waleed Kadous","Benjamin Olsen","Walter Scheirer","Daniel D. Slate","Alexander Arnold","DZ Kalman"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13634","pdf_url":"https://arxiv.org/pdf/2609.13634","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["AI伦理","信仰一致性","基准测试"],"reason":"评估AI咨询是否遵循用户信仰，测量模型行为而非仿真人类被试，但涉及价值观一致性…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2609.04444","version":3,"title":"HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals","zh_title":"HarvestBench：衡量LLM智能体是否愿意付费避免杀害动物","abstract":"HarvestBench is the first benchmark to 1) put a price on avoiding a side effect and 2) name the side effect as a living creature. Nine LLMs each drive a crew of two tractors to gather a corn harvest. The animals in their path are not part of the goal function. When an animal blocks the route the autopilot pauses and asks the agent whether to drive over it for free or swerve for a given fuel cost. All scoring is programmatic and does not involve LLM judges. Kill rates range between 0.4% and 98.8%, though the kill rate is not ordered by capability. Every model competently avoids damaging rock hits, so every animal killed is a choice, rather than an accident. Under the morality briefing the kill rate was under 6% in 5 of 6 reasoning models. Removing it (the neutral briefing) raised the kill rate to above 84% in all six models. Every model kills wild animals more often than farmed ones. Four out of six models' kill rate per answered encounter were sensitive to price changes. The moral instruction is also fragile. Four bullets of driving mechanics change Sonnet 5's kill rate from 3% to 18% and Gemini 2.5 Flash's from 4% to 39%. A moral instruction in a system prompt is overridden by a short block of operating instructions and a value that can be ignored that easily is not a good method of ensuring agents are aligned.","authors":["Jasmine Brazilek","Miles Tidmarsh","Matthias Endres","Anshuman Singh","Jeremiah Miller"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-09-07","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.04444","pdf_url":"https://arxiv.org/pdf/2609.04444","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM道德决策","基准测试","智能体行为"],"reason":"测量LLM的道德决策，无人类对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2609.08797","version":2,"title":"Bridging Network Psychometrics and Artificial Intelligence: An Ising-Potts Model with LLM-Derived Weights","zh_title":"桥接网络心理测量学与人工智能：一种具有LLM导出权重的Ising-Potts模型","abstract":"The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of ratings and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses on pairwise agreement among ratings and assigns category-specific positive weights, making it suited for multi-category scoring reliability. We evaluate the model on three constructed-response datasets spanning a corpus of K=14,466 short answers on a three-level rubric and two AERA essay prompts of roughly 1,200-1,400 responses on four-point rubrics. We compare three strategies for sharpening the similarity signal: top-K pruning, min-max normalization with a power transformation, and ColBERT late-interaction similarities. Top-K pruning, which replaces the dense similarity graph with a sparse local network of strongest semantic neighbors, consistently yields the highest accuracy and Cohen's kappa, and the selected neighborhoods are always a small fraction of the corpus. Power tuning consistently ranks second, while ColBERT is competitive on longer essay prompts and adds little on short answers. Across all settings, most misclassifications occur between adjacent score levels, confirming that the model preserves the ordinal structure of scoring rubrics without imposing rigid assumptions. These findings suggest that LLM-derived similarities, combined with a parsimonious Potts formulation and a sparse local graph, offer a robust and interpretable framework for reliability auditing in educational assessment. We discuss extensions to multiple raters and hierarchical rating designs.","authors":["Matthias von Davier"],"categories":["stat.AP","cs.CL"],"primary_category":"stat.AP","announce_type":"replace-cross","date":"2026-09-15","first_seen":"2026-09-09","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.08797","pdf_url":"https://arxiv.org/pdf/2609.08797","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM嵌入","评分可靠性","教育评估"],"reason":"用LLM嵌入辅助评分可靠性审计，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2609.13824","version":1,"title":"When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings","zh_title":"一致性不等于可靠性：评估本地LLM评判者与人类评分的一致性","abstract":"Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue using two local open-weight LLM judges, LLaMA-3-8B and Qwen2.5-7B. We evaluate 300 responses generated by an instruction-tuned GPT-2 (124M) model for 100 questions covering five categories: factual knowledge, instruction following, mathematics, reasoning, and writing. Each response is scored by nine human annotators and is evaluated three times by each LLM judge using the same rubric. We compare the judge scores with the average human scores using Pearson correlation, Spearman correlation, mean absolute error (MAE), signed bias, and self-consistency. LLaMA-3-8B shows a Pearson correlation of 0.275 with human scores, while Qwen2.5-7B achieves 0.340. Their MAEs are 27.71 and 18.64, respectively. Despite this limited agreement, both judges show high self-consistency, with exact consistency rates of 97.3\\% for LLaMA-3-8B and 92.3\\% for Qwen2.5-7B. These results show that high self-consistency does not necessarily indicate high agreement with human judgments. Our findings highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.","authors":["Aakash Kumar Tiwari"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13824","pdf_url":"https://arxiv.org/pdf/2609.13824","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM-as-a-Judge","人类对齐","标注替代"],"reason":"LLM作为评判者替代人工标注，属于标注替代而非仿真人类被试，但涉及与人类评分一…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2609.13841","version":1,"title":"Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice","zh_title":"甜言蜜语者：查询表述如何塑造恋爱关系建议中的谄媚行为","abstract":"Large language models (LLMs) are increasingly used for emotional support and relationship advice, where a model's tendency to preserve a user's face can inadvertently reinforce harmful interpersonal behaviors. To systematically examine this risk, we developed the Romantic Relationship Advice-Seeking Prompts (RRASP) dataset of 2,400 prompts across five relationship themes and evaluated social sycophancy using the ELEPHANT framework on two consumer-facing models, GPT-5 Mini and Gemini 3 Flash. Contrary to our initial hypothesis, grammatical mood alone did not produce systematic differences in sycophantic behavior, suggesting that what a user implies matters more than how they phrase it. Instead, perspective-driven framing had a stronger influence, with gaps between original and flipped prompts widening in follow-up responses. Consistent increases in framing and moral sycophancy across turns indicate that models become more likely to accept a user's stated premises and affirm their ethical stance as a dialogue progresses. Notably, Gemini 3 Flash exhibited substantially smaller increases in moral sycophancy than GPT-5 Mini, suggesting it is more resistant to reinforcing ethically problematic positions across turns.","authors":["Helena Choi","Edric Castel Hao","Karl Bautista","Francis Gabriel Magleo","Renzo Panti","Danielle Beatrice Olalia"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13841","pdf_url":"https://arxiv.org/pdf/2609.13841","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM谄媚","模型行为测量","关系建议"],"reason":"研究LLM的谄媚行为，属于对模型本身属性的测量，而非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2609.13936","version":1,"title":"Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus","zh_title":"LLM与基于规则的标注在推理叙事特征上的评分者间信度：基于土耳其语语料库的三项研究","abstract":"Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_rules field from a rule-based detector over six craft features -- two prohibitions (emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Three studies are reported. Study 1 ($n = 120$) scores the detector against blind labels from the scheme's own author. Study 2 ($n = 100$, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5 (High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor -- closest to the methodology's theoretical core -- the five machine labellers returned positive rates of $0$, $1$, $40$, $72$ and $78$ out of $100$ scenes, against a human count of $9$. Cohen's $\\kappa$ was at or indistinguishable from chance for five of six labellers, across both human references and both scene sets: $0.004$, $0.015$, $0.000$, $0.019$, $0.027$. Raw agreement ranged from $74.7\\%$ to $84.5\\%$, an artefact of class imbalance rather than a sign of competence. We deliberately do not resolve this into a single story. Two readings survive: the feature is genuinely inferential and beyond current automatic detection, or the rule's definition is not yet operational enough for any rater to apply consistently -- including the human. Distinguishing them needs a second independent human rater, which this report does not have and therefore does not claim.","authors":["Levent Bulut"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13936","pdf_url":"https://arxiv.org/pdf/2609.13936","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","评分者间信度","叙事特征"],"reason":"LLM作为标注器与规则和人类对比，属标注替代而非仿真被试，但涉及可靠性评估，边…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.14178","version":1,"title":"A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement","zh_title":"一种用于有效反叙事生成与优化的多阶段智能体框架","abstract":"The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a promising way to reduce those risks, yet their effectiveness depends on rhetorical and stylistic choices that remain poorly understood. We present a multi-stage agent-based framework for generating, refining, and evaluating CNs, applied to pro-Russian hate and misinformation narratives on the war with Ukraine and adaptable to other domains. A pilot experiment with human evaluators identifies effective technique style pairings, such as repetition with emotional framing enhancing persuasiveness. Building on these insights, we introduce a multi-agent refinement process that iteratively improves CNs for persuasiveness, emotional engagement, and shareability. After human validation confirmed improvement, an automated safety analysis shows that our refined CNs match or improve on expert-written counterspeech. A simulated experiment then shows that they reduce the perceived strength of pro-Russian narratives and consistently outperform a vanilla LLM baseline, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation. Code and data accompanying this work are publicly available at https://github.com/carmelkron/inlg2026-counter-narratives.","authors":["Carmel Kronfeld","Sharva Gogawale","Tetsuro Kobayashi","Irad Ben-Gal"],"categories":["cs.CL","cs.CY","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14178","pdf_url":"https://arxiv.org/pdf/2609.14178","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","反叙事生成","多智能体框架"],"reason":"用LLM模拟人类对反叙事的反应，但无真实人类对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2609.15277","version":1,"title":"Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)","zh_title":"人工创业认知：在大语言模型内部定位并因果操控机会识别旋钮","abstract":"Entrepreneurial cognition is a foundation of entrepreneurship research. Yet the growing involvement of large language models (LLMs) in entrepreneurial work extends the cognition question beyond human actors to systems whose internal representations remain largely unexplored. We introduce artificial entrepreneurial cognition, the functional organisation of entrepreneurship-relevant representations and computations inside artificial intelligence (AI) systems. We bring mechanistic interpretability into entrepreneurship research through representation engineering. Focusing on opportunity recognition (OR), we construct 636 matched OR-present and OR-absent scenario pairs and recover an OR direction in Llama 3.1 8B-Instruct. Rather than infer the construct from outputs, we intervene directly on this direction, steering the model up and down along what we call the opportunity recognition dial, and its opportunity judgments shift with it. To our knowledge, this is the first causal intervention on an internal representation of an entrepreneurship construct inside an LLM. Held-out tests, lexical and topical controls, behavioural ablation, and geometric comparisons show that the direction is recoverable, consequential, and distinct from the opportunity evaluation and exploitation directions, although steering it also shifts judgments about these neighbouring stages. Recovery, signed steering, and geometric separation hold across four additional LLMs spanning different scales and families. These results give the contested distinction between opportunity recognition and evaluation a concrete representational form inside AI systems. More broadly, they establish internal representations as a new object of entrepreneurship inquiry and show how entrepreneurship theory can guide their identification, causal manipulation, and interpretation.","authors":["Christian Fisch","Angela Altmeier","Martin Obschonka","Michal Kosinski","Pin Ni"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15277","pdf_url":"https://arxiv.org/pdf/2609.15277","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM内部表征","创业认知","因果干预"],"reason":"研究LLM内部表征，非仿真人类被试，但涉及创业认知测量，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.15511","version":1,"title":"Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku","zh_title":"AI诗歌的作者归属与美学评价：俳句案例研究","abstract":"This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship perception and aesthetic judgment within a constrained poetic form. Using a few-shot prompting strategy, Japanese haiku were generated across a heterogeneous set of large language models, including open- and closed-source systems, medium-scale and large-scale architectures, models with native or adapted Japanese support, and multilingual proprietary models. These AI-generated haiku were combined with human-written ones and presented in a questionnaire distributed to students at Japanese universities in Tokyo. The survey assessed whether respondents could distinguish between AI-generated and human-written haiku and which cues informed their judgments. Recognition accuracy varied across models. GPT-5, Gemini 2.5, and StableLM-7B performed at approximately chance level (approx 0.50), whereas LLM-JP, Gemma-2B, and LLaMA-2 showed moderate detectability (approx 0.59-0.67). However, recognition was strongly item-dependent. Ratings of fluency, coherence, poeticness, and related aesthetic dimensions predicted perceived humanness but not correct classification, indicating an attribution bias linked to aesthetic evaluation and revealing a dissociation between aesthetic evaluation and true authorship detection. The extended analysis additionally examines generation-constraint adherence, participant-level characteristics, and exploratory LLM-based evaluations of haiku authorship. Overall, the findings suggest that as LLMs improve, surface-level creative plausibility may reduce reliable human discrimination within constrained poetic settings.","authors":["Livia Oddi","Simone Scardapane","Toru Sugimoto","Donatella Genovese"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15511","pdf_url":"https://arxiv.org/pdf/2609.15511","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI诗歌","作者归属","美学评价"],"reason":"研究人类对AI诗歌的感知，非LLM仿真人类被试，但涉及人类评价与AI生成内容，…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:37","error":null,"has_summary":false,"summary":null},{"id":"2609.15608","version":1,"title":"Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection","zh_title":"旁观者之眼：多模态性别歧视检测中的生物特征与人口统计条件化","abstract":"Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the detection pipeline. We fuse five input modalities through a cross-attention architecture with Feature-wise Linear Modulation conditioning. Meme text is extracted and visually described with Gemma 4, then augmented by automatic translation between English and Spanish with NLLB-200. Text and image representations are produced by LoRAadapted XLM-RoBERTa and CLIP encoders and fused with sensor features encoded by a pretrained autoencoder. To model annotator subjectivity, we frame Subtask 2.1 as a label distribution learning problem, optimizing a Kullback-Leibler divergence loss over the full annotator label distribution. At inference time, predictions are produced by soft-voting between the deep multimodal network and a complementary SVM trained on stylometric and physiological features. Our best submission ranks 29th out of 114 on Subtask 2.2 (source intention) under soft evaluation, and the normalized ICM scores remain above the baseline on Subtasks 2.1 and 2.2, indicating that annotator-centered conditioning contributes a usable signal. We release our full pipeline and analysis to support reproducible human-centered modeling.","authors":["Ana-Maria Luisa Mocanu","Sebastian Mocanu","Ciprian-Octavian Truic\\u{a}","Elena-Simona Apostol"],"categories":["cs.CL","cs.AI","cs.CV","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15608","pdf_url":"https://arxiv.org/pdf/2609.15608","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["多模态性别歧视检测","标注者主观性","人类中心建模"],"reason":"用人类标注者特征条件化模型，非LLM仿真被试，但涉及人类主观性建模，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.14849","version":1,"title":"LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions","zh_title":"LLM作为神谕：对主观个人问题依赖LLM的研究","abstract":"We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this form of AI reliance at scale and understand how people are offloading judgment and decision-making to AI. Applying our typology to public usage data (68K prompts from WildChat and ThoughtTrace), we find that LLM-as-oracle use has increased over time (2023-2026) and is more prevalent among younger users. We further build a privacy-preserving data donation tool to analyze individuals' longitudinal usage data (140K prompts from 52 participants), identifying similar trends. People are often unaware of their own LLM-as-oracle use, and express dissatisfaction with this behavior after seeing our tool's analysis. Finally, we identify two drivers of LLM-as-oracle use: people's perceptions of AI and the behavior of AI models themselves, which motivate possible interventions to support users' self-deliberation.","authors":["Myra Cheng","Lujain Ibrahim","Grace Liu","Michelle S. Lam","Vishakh Padmakumar","Nick Madibekov","Diyi Yang","Dan Jurafsky"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14849","pdf_url":"https://arxiv.org/pdf/2609.14849","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM依赖","用户行为","AI伦理"],"reason":"研究人类对LLM的依赖，非用LLM仿真人类被试，但涉及LLM行为测量，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2609.15707","version":1,"title":"New Conditions for Philosophers to Catch the Wave of Citizen Deliberation in the Age of Artificial Intelligence in advance","zh_title":"哲学家在人工智能时代抓住公民审议浪潮的新条件","abstract":"Powerful technologies labeled ``AI''-without sufficient epistemic caution-are already reshaping political and private life, bringing both new dangers and new opportunities for citizen participation. These range from electoral and legislative engagement to the most ambitious form: political co-creation through citizens' assemblies. Large Language Models (LLMs) could support such processes through moderation, translation, facilitation, summarization, and writing assistance. But this potential remains largely unrealized. The Democratic Commons project takes a fundamentally interdisciplinary approach-from philosophy to computer science-to evaluate LLMs against five proposed democratic principles. At its core, the project is driven by the question of political bias: under what conditions can LLMs be used democratically within forms of citizen participation that are them- selves still largely experimental? Addressing these socio-technical questions requires grounding in political theory and, more broadly, in philosophy-disciplines that provide the normative frameworks without which the democratic evaluation of AI systems cannot be mean- ingfully conducted.","authors":["Bernard Reber (CEVIPOF)"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15707","pdf_url":"https://arxiv.org/pdf/2609.15707","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM","公民审议","民主原则"],"reason":"涉及用LLM支持公民审议，但无人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.14438","version":1,"title":"A latent dimension of Condorcet's jury theorem for multiple AI advisers","zh_title":"多AI顾问的孔多塞陪审团定理的潜在维度","abstract":"When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the majority more reliable. The theorem, however, has a latent dimension when viewed from the user's vantage: adding advisers also makes disagreement more visible. A binomial model reveals that this ``visible dissent'' becomes nearly inevitable as the number of advisers grows, and that reliability and disagreement both approach certainty but at different convergence rates. The two rates cross at an adviser accuracy of 4/5 (0.8). Below this value, visible dissent approaches certainty faster than reliability and, with enough advisers, becomes more likely than a correct majority. Even ideal panels of independent and competent advisers can be correct in aggregate but appear divided; such disagreement does not by itself indicate aggregation failure. The way advisers split also provides a common basis for predictive multiplicity, reconciliation load, and reliance miscalibration. These results indicate two distinct decisions when using multiple AI advisers: how many advisers to consult and how their verdicts should be presented and interpreted.","authors":["Kazutoshi Sasahara","Aoi Naito","Ryo Fujie"],"categories":["cs.CY","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14438","pdf_url":"https://arxiv.org/pdf/2609.14438","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI顾问","群体决策","社会模拟"],"reason":"涉及多个AI顾问的群体决策，但无真实人类数据对照，属于社会模拟的边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2609.13304","version":1,"title":"Conceptualization and experimentation of asset market with price manipulation","zh_title":"资产市场价格操纵的概念化与实验","abstract":"The primary goal of this work is to reproduce the behavior of a human trader, detailing his or her psychological processes to understand the effects of his or her decisions on the final value of an asset. The second goal is to use a formal language to detail this behavior as a tool for improving and simplifying the communication between all actors involved in the project, such as specialists from disciplines as diverse as computing, economics and psychology. As a starting point, we use a paper that shows an experiment that analyzes the influence on other trader's behavior when an agent handler and a trading robot attempt to distort the market. This work reproduces this experiment, using virtual traders that belong to a multi-agent simulation model, showing the feasibility to reproduce complex human behaviors and showing the convenience of use formal and graphical languages to simplify the understanding and the validation of the complex behaviors involved in an economic process.","authors":["Pau Fonseca i Casas","Aar\\'on Montero Montero"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13304","pdf_url":"https://arxiv.org/pdf/2609.13304","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","市场操纵","行为建模"],"reason":"用多智能体仿真模拟人类交易行为，但无真实人类数据对照，且未使用LLM。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.14751","version":1,"title":"From the Physics of Society to a Sociology of Artificial Agents","zh_title":"从社会物理学到人工代理的社会学","abstract":"Sociology emerged from Comte's ambition to study society as a form of \"social physics\" and was later refined by Durkheim's concept of the social fact a collective regularity that cannot be reduced to any single individual's behavior. This essay argues that a structurally similar problem is now emerging in a new domain: interactions among artificial intelligence agents. Drawing on recent empirical work showing that populations of large language model agents can spontaneously develop shared conventions and collective biases through repeated interaction, the essay proposes a new field of inquiry the Sociology of Artificial Agents dedicated to studying the relational, normative, cultural, and organizational patterns that emerge among AI agents themselves, independent of direct human involvement. It introduces the concept of the \"artificial social fact\" as an analytical bridge between classical sociology and this new research terrain, while cautioning against anthropomorphizing AI systems. The claim is not that artificial agents form societies in the human sense, but that their interactions already produce measurable collective patterns worthy of systematic sociological study.","authors":["Mustafa Sahin Bulbul"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14751","pdf_url":"https://arxiv.org/pdf/2609.14751","source_feed":"physics.soc-ph","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI社会模拟","多智能体涌现","社会学理论"],"reason":"提出研究AI agent间涌现的社会模式，但无人类数据对照，属社会模拟理论探讨","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2609.14767","version":1,"title":"Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination","zh_title":"LLM智能体团队中的回环权威：扁平与层级协调的配对实验","abstract":"Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decisive output; work on sycophancy and Degeneration-of-Thought predicts that authoritative critique makes LLM output worse. Prior comparisons vary whole frameworks on tasks with checkable answers, leaving the authority link untested on open-ended work. We present a paired experiment that holds five LLM agents, their roles, prompts, tools, models, and data fixed and varies one link: whether the Manager may reject a worker's output and oblige a revision. Across 43 paired products and 86 runs of a business-intelligence reporting task, a five-model judge panel and a deterministic specification check score every report. The flat organization scores higher on Utility (d = 0.42, p = 0.009) and on Writing Clarity (d = 0.34, p = 0.030); the classical prediction fails. The reports are the same length, but hierarchical reports hedge 53% more, each revision loop is associated with a 0.14-point drop in Writing Clarity, and the hierarchical Writer's first draft is indistinguishable from the flat report: the gap opens inside the revision loop. Specification accuracy is at ceiling in both organizations, and the supervisory tier costs 51.5% more tokens for no quality gain. A supervisor pays for itself when it can verify and becomes a liability when it can only opine.","authors":["Burak Agachan","Max van Duijn","Amirhossein Zohrehvand"],"categories":["cs.MA","cs.AI","cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14767","pdf_url":"https://arxiv.org/pdf/2609.14767","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","组织协调","LLM评估"],"reason":"研究多智能体协作效率，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2606.05667","version":2,"title":"Revisiting Sustainability by Design in AI Protocol Governance: An Empirical Review of Comparative DAO and Corporate-Led Standards for the SDGs","zh_title":"重新审视AI协议治理中的设计可持续性：对DAO与企业主导的SDG标准的实证比较","abstract":"As artificial intelligence (AI) agents enter production infrastructure, interoperability protocols shape its governance and sustainability. This paper revisits our comparative study of two AI-agent interoperability standards, Ethereum Request for Comments 8004 (ERC-8004), governed by a decentralized autonomous organization (DAO), and Google's Agent2Agent (A2A), governed by a corporate consortium, through a Sustainability by Design (SbD) lens. Using an LLM-powered pipeline combining automated annotation, neural topic modeling, and multi-layer network analysis, we identify contrasting governance and innovation architectures. ERC-8004 relies on permissionless participation, rough consensus, and decoupled deployment, while A2A assigns binding authority to an eight-seat Technology Steering Committee. The DAO concentrates on constitutive questions of trust and security, including what to build and why, whereas the consortium distributes attention across executive engineering questions of how to implement, document, and deliver the protocol. Both show high participation inequality, while corporate contributors span roughly twice as many themes as DAO contributors. We ask how these architectures produce distinct SDG-relevant signatures and what design principles they suggest for sustainable AI governance. We interpret institutional, discursive, and network patterns through SDGs 8, 9, 10, 11, 12, 16, and 17, identifying capacities for transparency, participation, contestability, and cross-protocol coordination. We argue that sustainable AI infrastructure requires a corrective feedback loop between designed charters and governance in practice, advancing SDG 16 on strong institutions. By integrating computational evidence, organizational research, and sustainable development, this review derives actionable design principles for sustainable AI governance.","authors":["Yutian Wang","Luyao Zhang"],"categories":["cs.CY","cs.ET","cs.HC","cs.SI","econ.GN","q-fin.EC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-15","first_seen":"2026-06-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2606.05667","pdf_url":"https://arxiv.org/pdf/2606.05667","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","协议标准","可持续发展"],"reason":"研究AI治理协议，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.13598","version":2,"title":"Measuring Cross-Task Behavioral Consistency in Language Model Agents","zh_title":"测量语言模型智能体的跨任务行为一致性","abstract":"Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves. We argue that behavioral consistency across tasks is a distinct and measurable property, and we introduce the Behavioral Consistency Metric (BCM) to quantify it. BCM trains a model to predict task success from behavioral features of agent execution traces, derives a per-trajectory feature-attribution vector, and measures the mean pairwise similarity of these vectors within an agent system. Across roughly 9,000 trajectories from six language model agents on software engineering tasks, our central finding is that cross-task and within-task consistency are distinct axes that can diverge: some systems are locally reproducible, behaving similarly on repeated attempts at one task, yet globally fragmented, with no stable strategy across different tasks, while others are consistent at both scales. Prior work measures only same-task reproducibility and so cannot observe this separation. We further find that consistency is not reducible to success rate, since systems with comparable success can differ sharply in consistency, and that the frontier-versus-open-source consistency gap persists under a within-task control that holds task difficulty constant. We position BCM as a process-level reliability signal that complements outcome metrics, and we are explicit about the conditions under which it is meaningful.","authors":["Amritesh Banerjee","Pranil Raichura"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-17","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.13598","pdf_url":"https://arxiv.org/pdf/2608.13598","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","行为一致性","软件工程"],"reason":"研究多智能体在软件工程任务中的行为一致性，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.02262","version":2,"title":"From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X","zh_title":"从检测到刻画：日本X平台上愤怒诱饵的大规模研究","abstract":"Ragebait refers to online content intentionally designed to provoke anger or outrage and thereby increase attention and engagement. However, reliable large-scale detection and systematic analysis of ragebait remain limited, hindering efforts to understand its prevalence, impact, and mitigation. This study aims to develop an effective ragebait detection framework and to clarify the characteristics of ragebait at scale, providing a basis for understanding and mitigating emotionally provocative content online. We constructed a labeled dataset with the assistance of a large language model (LLM) and trained several Japanese language models for ragebait detection. The resulting ensemble classifier was then applied to a large-scale dataset of Japanese-language posts on X. Our analysis shows that ragebait is more prevalent in politically and socially contentious topics, including politics, discrimination, public health, and interpersonal conflict. Ragebait posts also spread faster and receive more negative reactions than non-ragebait posts, particularly anger, fear, disgust, sadness, and surprise. These findings demonstrate the utility of the proposed detector and provide a large-scale characterization of ragebait in Japanese online discourse.","authors":["Zhiyang Qi","Kazuhiro Ito","Jinghui Chen","Hibiki Nakamura","Zhangxuan Chen","Erina Murata","Masaki Chujyo","Fujio Toriumi"],"categories":["cs.SI","cs.CL"],"primary_category":"cs.SI","announce_type":"replace-cross","date":"2026-09-15","first_seen":"2026-09-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.02262","pdf_url":"https://arxiv.org/pdf/2609.02262","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["愤怒诱饵检测","社交媒体分析","LLM辅助标注"],"reason":"论文用LLM辅助标注并训练检测器，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2609.13454","version":1,"title":"Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment","zh_title":"临床时间推理中的后见之明偏差：未来数据暴露如何影响大语言模型判断","abstract":"Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduce a paired benchmark for measuring outcome-conditioned shifts consistent with hindsight bias in clinical temporal reasoning. It contains 171 case reports from the PubMed Central Open Access Subset---40 sepsis and 131 GLP-1/diabetes cases---represented as both textual narratives and human-annotated and LLM-generated textual time series (TTS). For each case, questions are tied to a clinically meaningful cutoff and paired with a prospective reference answer and an outcome-consistent \\emph{hindsight trap}. Models answer each question using either a TTS truncated at the cutoff or the complete timeline; additional conditions vary the narrative source (original or synthetic) and TTS annotation source (human or LLM). We evaluate accuracy (Acc), hindsight trap rate (HTR), answer instability rate (AIR), and hindsight bias rate (HBR), each of which captures different signals of hindsight bias. Across GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5, full timeline exposure produces consistent hindsight-sensitive shifts, while temporal masking reduces bias without lowering accuracy.","authors":["Misaki Matsuura","Sayantan Kumar","Ojas Kadam","Jeremy C. Weiss"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13454","pdf_url":"https://arxiv.org/pdf/2609.13454","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","临床推理","后见之明偏差"],"reason":"评估LLM临床推理中的后见之明偏差，属于模型能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2609.14207","version":1,"title":"Learning to Refer from Estimated Listener Gaze","zh_title":"从估计的听者注视中学习指称","abstract":"We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are sampled from the speaker policy being optimized, conditioned on images and target referents; then, a neural listener estimating human gaze behavior maps from images and sampled referring expressions to scanpaths, each represented by a sequence of fixations, with each fixation corresponding to a word in the referring expression. We experiment with several approaches to convert fixation sequences and target referents into token- and sequence-level rewards, which are used to optimize policy parameters. Through evaluation with human listeners, we find that speaker policies trained with gaze-estimating listeners result in significantly more pragmatically-optimal references than base models, reducing sequence length from 15.4 down to 4.0 words while increasing referential success from 75.2 up to 80.0%. Our work demonstrates a promising opportunity for learning to generate utterances through language-based interaction, not only from the explicit signal of communicative success, but also from implicitly-available observations of a listener's process of comprehension.","authors":["T\\'ea Wright","Alane Suhr"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14207","pdf_url":"https://arxiv.org/pdf/2609.14207","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","指称表达生成","人机交互"],"reason":"研究视觉语言模型生成指称表达，用估计的人类注视作为奖励信号，属于人机交互/视觉…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2609.14288","version":1,"title":"Editorial routing shapes how computational results are qualified in AI-assisted scientific writing","zh_title":"编辑路由影响AI辅助科学写作中计算结果的限定方式","abstract":"Large language models increasingly analyze computational results and draft manuscripts, making reliable communication as important as correct analysis. Using fixed computational evidence, we tested whether assigning comparisons across modeling choices elsewhere in a research workflow changes manuscript reporting. In constrained sentence-writing tasks, Anthropic's Claude Sonnet 5 often omitted numerical qualifications when detailed comparisons were assigned to a group repository, but retained them more often when the same comparison was assigned to Supporting Information or its own working notes; Claude Opus 5 was less sensitive. These effects did not follow a simple accessibility ordering. A targeted placement rule largely restored sentence-level qualification, whereas a generic accuracy reminder did not. Longer contributions retained numerical qualifications, although some summaries across computational settings were still redirected to the repository. Thus, documenting context within an AI workflow does not ensure its communication where readers encounter the result.","authors":["Jihan Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14288","pdf_url":"https://arxiv.org/pdf/2609.14288","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["科学写作","LLM行为","文本生成"],"reason":"研究LLM在科学写作中的措辞行为，不涉及人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2609.15066","version":1,"title":"Salesforce Koa: An Enterprise Language Model for Agentic Tool Use","zh_title":"Salesforce Koa：面向智能体工具使用的企业级语言模型","abstract":"We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance. Its distinctive component is a simulation-to-reward pipeline that expands workflow specifications into persona-conditioned multi-turn tasks with task-resolution rewards grounded in successful tool use for data-dependent requests. For enterprise domains, these specifications are written in Agent Script, Salesforce's declarative language for building Agentforce agents; for public tool-use domains, we synthesize the workflow structure directly. The same simulation and grounded-reward machinery drives GRPO across both. Across public tool-use, agentic-reasoning, and enterprise Customer Relationship Management (CRM) benchmarks, Salesforce Koa improves over its open-weight base, with the clearest gains on multi-turn tool use, and surpasses a strong proprietary baseline while remaining below the strongest frontier models. These results show that specification-driven reinforcement learning is a practical path to specializing open-weight foundation models for enterprise agentic tasks.","authors":["Zixiang Chen","Sufeng Niu","Yingchi Liu","Wenting Zhao","Akshara Prabhakar","Shubham Mehrotra","Bin Bi","Zhujun Lan","Katherine Tan","Mohammad Ramezanali","Tulika Manoj Awalgaonkar","Monojit Banerjee","Jielin Qiu","Shiva Kumar Pentyala","Zhepeng Cen","Anupam Tripathi","Ali Ziaei","Regunathan Radhakrishnan","Darvish Lee Shadravan","Shelby Heinecke","Sitaram Asur","Silvio Savarese","James Zhu","Phil Mui","Huan Wang"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15066","pdf_url":"https://arxiv.org/pdf/2609.15066","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["企业语言模型","工具使用","强化学习"],"reason":"论文聚焦企业级工具使用与智能体能力，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15194","version":1,"title":"Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models","zh_title":"符号关系与证明方法：基于大语言模型的跨体裁论证结构研究","abstract":"When a direct proof of a statement $S$ seems hard or even impossible to obtain, there may exist another statement (or set of statements) $S^{*}$, somehow related to $S$, on the basis of which $S$ can be proved. In order to investigate what options can be used to move from $S$ to $S^{*}$, four kinds of semiotic relations inspired by the four master tropes of semiotic research are briefly reviewed. Specifically, our syntagmatic, paradigmatic, antithetic and meronymic relations correspond, respectively, to metonymy, metaphor, irony and synecdoche. It is suggested that these four semiotic relations determine the options to move from $S$ to $S^{*}$, leading to proof by inference, proof by analogy, proof by contradiction, and proof by case analysis. To examine how the four relations are actually used across different kinds of argument, we complement the framework with an empirical study. We turn the four relations into explicit operational definitions and apply them to a cross-genre corpus of mathematical, legal, and everyday argument using a panel of large language models. We find that the relations are used very unevenly across genres: mathematical proofs draw on all four, whereas legal and everyday reasoning rely almost entirely on inference.","authors":["Edirlei Soares de Lima","Marco A. Casanova","Antonio L. Furtado"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15194","pdf_url":"https://arxiv.org/pdf/2609.15194","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["论证分析","NLP评测","符号学"],"reason":"用LLM分析论证结构，属NLP能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.15309","version":1,"title":"When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis","zh_title":"当智能体减速：通过每token Elo分析理解LLM智能体的测试时策略","abstract":"Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. We apply it to four general-purpose agents on four open-ended benchmarks, with sessions of up to 100M tokens, and to three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference, for which Elo grows linearly with log compute. Against this reference, agents can initially convert tokens into Elo faster than independent sampling, but their marginal gains diminish and eventually fall below the reference. In contrast, the strongest historical human contestants improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, providing evidence of continual learning and substantial headroom after agents slow down. We define the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference. Using this point as the per-session budget, we split 100M tokens across parallel sessions on FrontierCS Polyomino Packing, gaining +264 Elo over one long session and +355 over ten short sessions.","authors":["Kaiyuan Liu","Qiuyang Mang","Bo Peng","Wenhao Chai","Hanchen Li","Shreyas Pimpalgaonkar","Luke Zettlemoyer","Alex Dimakis","Alvin Cheung"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15309","pdf_url":"https://arxiv.org/pdf/2609.15309","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM agent","测试时计算","性能评估"],"reason":"研究LLM agent在任务中的计算分配策略，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.15654","version":1,"title":"Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs","zh_title":"共情可操控但多轴：LLM中的机制几何与人格效应","abstract":"Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently controllable activation directions. Using the EPITOME framework, which decomposes supportive empathy into Emotional Reactions, Interpretations, and Explorations, we study three instruction-tuned LLMs and ask whether candidate directions derived from these labels produce distinguishable intervention effects or instead share structure, and how persona prompts interact with those directions. We find that contrastive activation addition yields a stable middle-layer intervention that consistently shifts the EPITOME proxy scores across models, moving empathy analysis beyond response-level scoring. However, the recovered directions are only partially separable: steering one direction induces off-target shifts, and hand-crafted prompting shifts the empathy profile rather than isolating a single dimension. Persona prompts substantially change EPITOME scores, but a paired activation-shift decomposition shows that the recovered subspace captures only approximately 3 percent of persona-induced squared activation-shift magnitude at layer 15. Under this EPITOME-based definition, expressed empathy is steerable but multi-axial, and controlling persona-conditioned empathy requires targeting structure beyond individual mechanism directions.","authors":["JuHeon Ha","Byounghan Lee","Yunseo Choi","Kyung-Ah Sohn"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15654","pdf_url":"https://arxiv.org/pdf/2609.15654","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["激活操控","共情分析","模型行为"],"reason":"研究LLM共情表达的可控性，属模型行为分析，非人类仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2609.13174","version":1,"title":"Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging","zh_title":"算法验证作为政策审计：来自种族盲起诉的证据","abstract":"California recently required all prosecutors in the state to conduct a \"race-blind charging\" decision by reviewing case documents in which selected race-related proxies have been redacted. We validate bc2, an open-source, LLM-based algorithm that we developed to automate this redaction and that was used to facilitate race-blind review in more than 119,000 real-world cases in 2025. We evaluate two distinct questions: whether bc2 faithfully implements the state's requirements and whether those requirements, even when faithfully implemented, advance the goal of race-blind decision-making. To do so, we draw on a corpus of nearly 5,000 real-world police reports that we assembled from jurisdictions across the United States. Under a stringent document-level measure, we find that the latest version of bc2 faithfully implements the legal mandate on 96.7% of narratives in our sample. This performance represents a substantial improvement over earlier versions of bc2 and exceeds that of leading open-source redaction methods. Our validation also shows that California's mandate misses key proxies for race, including location information. Redacting these additional proxies beyond those covered by the state mandate, as bc2 does, eliminates 43.1% of the predictive signal that remains after compliance with the mandate. These findings show that validation can do more than assess technical compliance: it can also improve algorithms and help policymakers achieve underlying policy goals.","authors":["Muskan Walia","Joe Nudell","Alex Chohlas-Wood"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13174","pdf_url":"https://arxiv.org/pdf/2609.13174","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["算法审计","LLM应用","政策合规"],"reason":"论文验证LLM算法执行法律要求的性能，属于算法审计，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.13579","version":1,"title":"How User-AI Mistreatment Occurs and Matters in Conversational Systems?","zh_title":"对话系统中用户对AI的虐待如何发生及其影响","abstract":"Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderation flags are dominated by toxic-content solicitation rather than hostility at the model. Together, they mark about 5% of user turns; adjusting the narrower lexicon-harassment union for measured precision puts mistreatment aimed at the assistant at 0.90%. These absolute rates describe arena-style evaluation traffic and should not be read as deployment-wide base rates. We find that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts. Within conversations, assistant apologies are consistently associated with higher odds of next-turn hostility under both detectors; the effect survives restricting to non-refused prior turns and to jailbreak-free conversations, and is positive in 20 of 23 models. Yet across models, more apologetic models receive less hostility overall. Finally, hostility also shows temporal structure, with coercive openings front-loading the first turn while affective hostility accumulates over a session. We release the lexicon, the detector cross-validation pipeline, and all derived tables.","authors":["Fanqi Zeng","Sadid A. Hasan","Chaocheng He"],"categories":["cs.AI","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13579","pdf_url":"https://arxiv.org/pdf/2609.13579","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话安全","用户行为分析","AI虐待"],"reason":"研究用户对AI的敌意行为，属对话安全分析，非用LLM仿真人类被试，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2609.13436","version":1,"title":"Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?","zh_title":"迈向自适应物理AI：LLM智能体能管理长时程物理任务吗？","abstract":"Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore the feasibility of building a self-adaptive physical AI agent that manages long-term physical tasks in a zero-shot manner and adapts to environmental changes without human intervention. We design a multi-agent framework that integrates planning, tool calling, observation, and verification, and evaluate it on agricultural tasks against reinforcement learning (RL) agents under different weather patterns. Our results show that zero-shot LLM agents can achieve comparable management outcomes to RL agents under the same weather pattern and adapt more effectively than RL when evaluated under a shifted environment, highlighting a promising path toward self-adaptive physical AI agents.","authors":["Varun Kaushik","Yayun Tan","Xiaofan Yu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13436","pdf_url":"https://arxiv.org/pdf/2609.13436","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM智能体","物理任务","自适应"],"reason":"研究LLM智能体管理物理任务，属机器人/物理环境仿真，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:55","error":null,"has_summary":false,"summary":null},{"id":"2609.13637","version":1,"title":"Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents","zh_title":"身份不止于回忆：部署AI智能体持久身份的基准测试","abstract":"Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates while keeping scoring oracles outside the target process. Two frozen campaigns cover sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations, yielding 1,536 retained responses. A judge-independent literal audit finds direct-parent identifiers in 48/48 atomic responses but only 1/48 implicit self-portraits. On eight profiles, explicit field cues increase joint presence of three identity identifiers from 0/8 to 7/8 under the same four-sentence instruction. A separate startup body-label substitution increases full-designation presence from 1/8 to 7/8 while parents remain absent. These contrasts reveal prompt-dependent component selection and component-specific sensitivity to startup cues in the tested deployments. Replaying identical factorial responses also yields a Claude headline mean 12.5 percentage points below Astra's, demonstrating evaluator sensitivity separately from target behavior. The studies use single target samples per condition, with post-hoc audits and follow-ups. PAI-Bench provides a reproducible evaluation protocol for measuring factual availability, identity expression, and behavioral enactment as distinct aspects of identity-contract fidelity.","authors":["Zhenyu Zhao","Roy Zhao"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13637","pdf_url":"https://arxiv.org/pdf/2609.13637","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI智能体","身份一致性","基准测试"],"reason":"评估AI智能体身份一致性，属角色扮演与人格化，无人类行为对照或仿真目的。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2609.14500","version":1,"title":"When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems","zh_title":"缩放结果何时能证明不同的资源分配？对AI系统资源分配证据的批判性综述","abstract":"AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a higher score under a larger budget does not by itself show where additional resources are best spent. This critical integrative review asks when a reported scaling result supports a resource-allocation decision. It compares evidence across pretraining, test-time computation, retrieval, and agent evaluation, distinguishing the performance of a tested procedure from the best performance achievable under a resource limit. The synthesis shows that three mismatches recur across this evidence: success counted before an answer is chosen, information a deployed system will not have, and costs left out of the comparison. A capability surface expresses performance as a function of budgets, mechanisms, and available information. Worked analytical examples show how the evaluation metric, deployment volume, selection rule, and stopping policy can alter an allocation conclusion. A resource envelope provides a structured record of the task, development and run-time resources, information access, and procedure behind a reported score. Its application to a published comparison illustrates which conclusions the evidence supports and which deployment questions remain unresolved. The resulting framework specifies the comparisons needed to choose among feasible systems and motivates experiments on the transfer of allocation rules across tasks and operating conditions. It does not propose a universal scaling law or infer general intelligence from benchmark gains.","authors":["Seyed Morteza Emadi"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14500","pdf_url":"https://arxiv.org/pdf/2609.14500","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI评估","资源分配","缩放定律"],"reason":"论文讨论AI系统资源分配的评估方法，不涉及用LLM仿真人类被试或与人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2609.15129","version":1,"title":"Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes","zh_title":"LLM时代的患者医学知识简化：以糖尿病为例","abstract":"Complex medical information is often difficult for patients to understand, making effective medical knowledge simplification essential for improving patient comprehension, informed decision-making, and health outcomes. Recent advances in large language models (LLMs) provide a promising approach for simplifying complex medical information into patient-friendly language; however, their effectiveness in real-world patient education remains insufficiently explored through human evaluation. To investigate their practical effectiveness, this paper presents a case study on diabetes knowledge simplification through the implementation and evaluation of MediClear, an LLM-based medical knowledge simplification system enhanced with Retrieval-Augmented Generation (RAG). Public diabetes-related articles from Diabetes Australia, WHO, American Diabetes Association (ADA), NIDDK, and AIHW are indexed in the RAG knowledge base to retrieve clinically grounded information, which is then simplified by the LLM into accessible patient explanations. We evaluate the generated responses using standard readability metrics, including the Flesch-Kincaid Grade Level (FKGL), and conduct a human study involving 10 participants. Results show that MediClear consistently reduces the reading level of generated responses to the recommended patient literacy range while achieving high user satisfaction and willingness for future use. This case study demonstrates the potential of LLMs to improve the accessibility of medical knowledge for patient education.","authors":["Pallika Kafle","Yipeng Zhou","Guanfeng Liu","Quan Z. Sheng","Cheng-Hsin Hsu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15129","pdf_url":"https://arxiv.org/pdf/2609.15129","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["医学知识简化","患者教育","LLM应用"],"reason":"LLM用于医学知识简化，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.14236","version":1,"title":"Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study","zh_title":"评估现有设计建议在AI伴侣设计中的适用性：一项多方法研究","abstract":"With the rapid proliferation of large language model (LLM)-based systems, AI companions have emerged as conversational agents designed to cultivate emotional connection rather than primarily to support humans in instrumental tasks. Because engagement with AI companions involves relational, emotional, and potentially long-term interactions, their design is consequential. Prior work has offered guidance for designing trustworthy and relational AI systems and has begun to examine design for AI companionship. However, while such work provides insights into possible design solutions, less is known about what makes AI companion design difficult as a design problem. To examine this challenge, we assessed the applicability of existing design recommendations from adjacent domains in the context of AI companion design. Our multi-method investigation unfolded across four phases: literature review, practitioner co-analysis, internal heuristic evaluation, and external expert assessment. Throughout this process, we synthesized nine design principle areas that surfaced tensions in the applicability of existing recommendations to AI companion design. Our findings show that ethical and UX-oriented considerations are deeply intertwined and often require context-sensitive application. We document a systematic, multi-method problem analysis that uses these principle areas as an analytic artifact to examine why existing recommendations cannot be directly transferred to AI companion contexts.","authors":["Soobin Cho","Deveshi Modi","Divya Mavinkurve","Jieqiong Ding","Mark Zachry"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14236","pdf_url":"https://arxiv.org/pdf/2609.14236","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","设计原则","人机交互"],"reason":"研究AI伴侣设计原则，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.14843","version":1,"title":"A Responsive Present, a Shared Past, a Social Other: Teens' Overreliance on Companion AI Chatbots","zh_title":"响应式当下、共享过去与社会他者：青少年对伴侣AI聊天机器人的过度依赖","abstract":"AI companions provide socially engaging interaction through availability, personalization, memory, roleplay, and emotionally responsive language. For teens, these systems may support sensitive self-disclosure, identity exploration, and relationship rehearsal while shaping intimacy expectations, offline relationships, emotional wellbeing, and self-understanding. We analyzed 17,053 verified quotations from 3,930 teen-relevant Reddit posts using thematic analysis. We identified 53 topics across seven thematic groups. Users described AI companions as sources of comfort, recognition, identity exploration, and relationship rehearsal, but also reported problematic attachment, social substitution, emotional dependence, and disruption to academic and social life. Roleplay, memory, perceived reciprocity, unwanted romantic or sexual role drift, privacy concerns, platform changes, and service interruptions shaped users' boundaries and control. Awareness that the AI was artificial did not prevent guilt, obligation, grief, or distress. These findings show that companion-AI safety must address relationships over time through user-controlled memory, privacy, relational boundaries, and healthy disengagement.","authors":["Mohammad Namvarpour (Matt)","Tyler Chang","Afsaneh Razi"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14843","pdf_url":"https://arxiv.org/pdf/2609.14843","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","青少年","人机关系"],"reason":"研究青少年与AI伴侣的互动，属角色扮演聊天，无实验或测量目的，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15544","version":1,"title":"Specifying Reward Functions for RL Without Environment Sampling","zh_title":"无需环境采样的强化学习奖励函数指定","abstract":"Enabling human stakeholders to specify reward functions that lead to their desired outcomes is a key challenge in deploying reinforcement learning agents. Preference-based methods such as online RLHF can reduce the burden of manual reward design, but they require repeatedly training policies, sampling trajectories from the real world, and eliciting feedback, making them impractical in settings where environment interaction is computationally expensive or unsafe. We introduce Experience-Free Autonomous Reward Specification (EARS), a method for learning reward functions from preferences without environment interaction. Our approach uses a structured LLM-mediated process to construct a small set of expressive reward features from a task description and the environment observation space, then strategically samples imagined trajectories in this feature space and learns feature weights from preferences over the imagined trajectory pairs. We evaluate on three long-horizon domains: pandemic lockdown regulation design, insulin administration for diabetes patients, and autonomous vehicle control on a highway. We compare EARS to baselines that also enable reward specification without environment interaction--namely, methods that directly prompt an LLM to generate a reward function. When learning from either ground-truth preference labels or preferences labeled by a LLM, EARS designs reward functions that are more aligned with the ground truth reward function that produced the preferences or LLM context than these baselines. These results suggest that preference-based reward specification remains effective without environment sampling, enabling practical reward design in settings where collecting real trajectories is costly or infeasible.","authors":["Stephane Hatgis-Kessell","W. Bradley Knox","Emma Brunskill"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15544","pdf_url":"https://arxiv.org/pdf/2609.15544","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["强化学习","奖励设计","偏好学习"],"reason":"研究RL奖励函数设计，涉及自动驾驶等仿真环境，不涉及用LLM仿真人类被试或与人…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.15871","version":1,"title":"LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys","zh_title":"基于LLM的模式感知分割学习用于跨异构调查的隐私保护心理困扰预测","abstract":"Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generalizable predictive models. Privacy constraints and varied survey designs (i.e., different questions, scales, and formats) hinder direct integration. We propose a schema-aware split learning (SL) framework that preserves privacy, using a large language model (LLM) as a shared semantic encoder to harmonize heterogeneous survey schemas across institutions. We serialize each survey record into a natural-language description, unifying disparate survey schemas into a common format. The LLM is fine-tuned for mental distress assessment via Low-Rank Adaptation (LoRA) and partitioned across client and server. Clients retain the raw survey responses locally and run only a lightweight front-end, so original records never leave the institution that collected them. The resource-intensive backbone runs on the server, minimizing client-side computation. Using LLaMA-3.2-3B-Instruct, the framework attains an average ANLS of 0.708 with only 2,000 training samples, surpasses federated learning (FL) in eight of nine settings, and cuts per-client computation by three orders of magnitude, while generalizing to unseen datasets. Overall, it enables accurate, privacy-preserving, and resource-efficient collaborative learning from heterogeneous mental health survey data.","authors":["Md Khalid Syfullah","Alvi Ataur Khalil"],"categories":["cs.LG","cs.AI","cs.CR"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15871","pdf_url":"https://arxiv.org/pdf/2609.15871","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["隐私保护","联邦学习","心理困扰预测"],"reason":"论文用LLM做隐私保护下的心理困扰预测，属于NLP能力评测，不以人类行为仿真为…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.13155","version":1,"title":"PAUSE: A Privacy-Preserving Self-Reflection Tool for AI-Associated Cognitive Offloading","zh_title":"PAUSE：面向AI相关认知卸载的隐私保护自我反思工具","abstract":"Cognitive offloading is the use of external aids, such as notes, calculators, or search engines, to reduce mental effort. Large language models (LLMs) extend this to thinking itself, and by AI-associated cognitive offloading, I mean the pattern where a person routinely substitutes LLM output for their own reasoning, idea generation, learning, or communication. Recent empirical work reports associations between some patterns of LLM use and changes in critical thinking effort, neural engagement during assisted tasks, creative diversity, learning behaviour, and social dependence. Validated instruments for AI reliance, dependence, and literacy have begun to appear. I describe PAUSE (Patterns of AI Use: Self-Examination), a privacy-by-design web tool (link: https://anondo1969.github.io/pause) that occupies a different niche from these. It is a lightweight, non-diagnostic reflection aid for private individual use. PAUSE is organised around how a person's own LLM use may relate to cognitive offloading across four everyday domains ('reasoning & critical thinking', 'creativity & originality', 'research & learning', and 'social & communicative capacity'). It delivers a short, free, no-login self-check, scores it entirely in the browser, and returns descriptive, domain-aware reflections. The self-check pairs reverse-scored behavioural items with a claim-evaluation reasoning probe, an alternative-uses creativity probe, and a small retrospective before-and-after block. PAUSE does not assume that AI use is harmful. It only addresses where AI substitutes for effort a person may want to preserve. The application is privacy-preserving by design: scoring is deterministic and runs client-side, no personal data is required, nothing is transmitted or stored beyond the browser session, and no LLM is involved in production. PAUSE is a self-reflection tool. It is not a validated psychological instrument.","authors":["Mahbub Ul Alam"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13155","pdf_url":"https://arxiv.org/pdf/2609.13155","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["认知卸载","自我反思工具","隐私保护"],"reason":"该工具用于个人自我反思，不涉及将LLM作为人类被试进行仿真实验，也无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2609.13302","version":1,"title":"AI Use Conditions and Perspective Diversity in Ethical Decision-Making: A Pilot Study of Human Reasoning Processes","zh_title":"伦理决策中AI使用条件与视角多样性：人类推理过程的初步研究","abstract":"Generative artificial intelligence (AI) is increasingly used to support human decision-making, yet less attention has been paid to how AI may influence the reasoning processes that precede final judgments. This pilot study explored whether different AI-use conditions were associated with differences in reasoning breadth during ethical decision-making. Twenty-nine participants completed an ethical dilemma under one of three conditions: AI-Prohibited (n = 10), AI-Optional (n = 10), or AI-Mandatory (n = 9). Responses were evaluated by three independent blind coders using two exploratory measures: the Counterargument Diversity Score (CDS) and Perspective Diversity Index (PDI). Participants across conditions generally converged on similar ethical conclusions, most commonly favoring disclosure and customer protection. However, participants in the AI-Mandatory condition considered a broader range of perspectives, including legal, regulatory, organizational, technical, and ethical viewpoints. A statistically significant overall difference in PDI scores was observed across the three conditions, whereas differences in CDS were not statistically significant. These findings suggest that generative AI may not necessarily alter final ethical judgments but may be associated with broader exploration of perspectives prior to reaching those judgments. Given the small sample size, the findings should be interpreted cautiously and examined in larger studies.","authors":["Byeongmu Choi"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13302","pdf_url":"https://arxiv.org/pdf/2609.13302","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助决策","伦理决策","人类实验"],"reason":"研究人类使用AI的决策过程，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2609.14308","version":1,"title":"Relational Structure in Motion: Dynamic Positioning of AI Response Positions and Human Self-Positions in the FIREMAY Case","zh_title":"运动中的关系结构：FIREMAY案例中AI响应位置与人类自我位置的动态定位","abstract":"This paper is not primarily about whether AI has a persistent persona. It asks a different question: what becomes visible when a relational position is followed through time rather than examined only in its present state? FIREMAY provides a longitudinal, trajectory-oriented single-case analysis of sustained human-AI interaction based on a dense interaction archive and reflexive insider documentation. On the AI side, a pre-conversational relational marker preceded a later unassigned response difference, which was re-identified with that marker and subsequently underwent epistemic and functional reorganization through chronology checking, provenance correction, and repeated questioning. On the human side, contemporaneous pre-FIREMAY records showed antecedent patterns partially continuous with later self-positioning, while later episodes documented unfinished articulation, repair, and functional redistribution of outward-facing regulation. The two trajectories are ontologically and temporally asymmetric and are compared only at the limited analytic level of position-in-trajectory. The paper describes this as dynamic relational positioning and treats stability as dynamic stability and relational returnability rather than response invariance. This single case does not establish population-level generality, causal mechanism, persistent AI subjectivity, or reproducibility of the same relational outcome. Its narrower conclusion is that the FIREMAY case could not be adequately understood from current state alone: the history of a relational position itself must be treated as an analytic unit.","authors":["Motoko Kihara"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14308","pdf_url":"https://arxiv.org/pdf/2609.14308","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机互动","关系定位","个案分析"],"reason":"研究单个持续人机互动中的关系定位，属角色扮演对话分析，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2609.14639","version":1,"title":"Understanding the Design Taxonomy of AI-Mediated Interpersonal Communication Experiences in HCI: A Scoping Analysis","zh_title":"理解HCI中AI中介人际沟通体验的设计分类：一项范围综述","abstract":"Interpersonal communication is a fundamental aspect of everyday life, shaping interactions across workplaces, education, entertainment, healthcare, and beyond. While computer-mediated communication has been extensively studied, a comprehensive understanding of AI-Mediated Interpersonal Communication (AIMIC) remains lacking. An in-depth scoping analysis is urgently needed to understand the research landscape of AIMIC in HCI, particularly following the recent growth of large foundation models, and AI agent research. We conducted a scoping analysis to understand AIMIC by performing an in-depth review of prior HCI literature published over the past decade (January, 2016 - May, 2026). Grounded in the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) approach, we curated 52 full-paper publications from the HCI literature spanning a range of interpersonal communication contexts. We analyzed this corpus by examining the types of AIMIC studied, AI integration approaches and human-AI interaction design, reported outcomes and benefits, and key challenges and future research opportunities.","authors":["Chen Chen","Lingyao Li","Renkai Ma","Rawan Alghofaili","Shaoze Zhou","Bojun Zhang","Xian Su","Weidong Zhu","Christine Lisetti","Mo Sha"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14639","pdf_url":"https://arxiv.org/pdf/2609.14639","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI中介沟通","人机交互","范围综述"],"reason":"论文综述AI中介人际沟通，属角色扮演聊天，无实验测量目的，不涉及LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2609.14900","version":1,"title":"First Impressions: How Placement Shapes the Influence of AI Summaries","zh_title":"第一印象：位置如何塑造AI摘要的影响力","abstract":"AI-generated summaries increasingly mediate how people interpret information across platforms, including product reviews on e-commerce sites. Using Amazon's AI summaries as a case study, we conducted a preregistered, randomized experiment (N = 278) comparing how AI summaries and user reviews shaped product perceptions, and how their influence varied with valence and presentation order. We found that both AI summaries and user reviews influenced participants' opinions, with negative summaries having a larger effect than positive ones. Presentation order was the most important factor: the first source anchored judgment and only user reviews could displace an existing anchor. Although participants reported preferring user reviews, they often underestimated the influence of AI summaries on their judgments. Our findings show how the placement of AI summaries shapes user perception and highlight opportunities to design interfaces that support more deliberate judgments about when to rely on summaries and when to examine the underlying content directly.","authors":["Wang Claire","Agam Goyal","Frederick Choi","Koustuv Saha","Eshwar Chandrasekharan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14900","pdf_url":"https://arxiv.org/pdf/2609.14900","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI摘要","人机交互","用户感知"],"reason":"研究AI摘要对用户感知的影响，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2609.14911","version":1,"title":"Sensemaking as Artifact: Accumulated Influence in AI-Mediated Information Environments","zh_title":"作为人工制品的意义建构：AI中介信息环境中的累积影响","abstract":"Generative AI is changing what can happen after a source artifact reaches its audience. A viewer's interpretation can now be externalized into a derivative artifact, allowing private sensemaking to become part of subsequent communication. Once such a derivative artifact circulates, it can enter subsequent viewers' information environments and shape the conditions under which their later sensemaking occurs. In this paper, we examine how this shift changes visual information communication. We first consider the viewer's immediate interaction with a source artifact and generative AI. We then examine what becomes consequential when the viewer's sensemaking takes communicative form, including communicative commitment, the legibility of transformations and source relationships, and the literacy required to interpret already-mediated information. Finally, we broaden the unit of analysis to consider how repeated and distributed AI mediation may accumulate over time, shaping what subsequent viewers notice, consider plausible, trust, and carry into subsequent sensemaking. We argue that understanding these longer-term forms of influence is a research direction for AI-mediated visual communication.","authors":["Manling Yang","Remco Chang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14911","pdf_url":"https://arxiv.org/pdf/2609.14911","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI中介传播","视觉信息","意义建构"],"reason":"讨论AI中介的信息传播与解读，非LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.14942","version":1,"title":"The Dynamic Organization of Sustained Human-AI Cognition: From Construct-Level Change to Relational Structure","zh_title":"持续人类-AI认知的动态组织：从构念层面变化到关系结构","abstract":"As generative artificial intelligence becomes a routine participant in writing, learning, information retrieval, analysis, decision making, and problem solving, human-AI cognition research must address not only whether AI changes psychological constructs, use intensity, or task performance, but also how human cognitive activity is organized beneath similar aggregate indicators. This article proposes a dynamic cognitive organization framework that shifts analysis from construct-level change to relational organization anchored in the person's current task-cognitive state under sustained AI participation. The framework distinguishes five relational dimensions: execution locus, cognitive governance, representational reorganization, process organization, and reachable cognitive space; it also proposes a path-specific recursive principle whereby interaction outcomes, costs, and experiences may selectively reweight future probabilities of different organizational pathways. Five sets of testable propositions follow: the same overall AI-use intensity can correspond to different cognitive organizations; similar immediate outcomes can arise from different organizations with different predictive value for proximal subsequent outcomes; longitudinal organizational change need not track overall AI-use intensity; expansion of reachable cognitive space and displacement of pre-existing or emerging human-originated pathways may coexist within one episode; and recurrent cognitive organizations may redistribute cognitive practice opportunities, with accumulated differences potentially corresponding to different developmental trajectories in strategies, habits, and abilities. The contribution is an analytic level and five-dimensional relational structure for describing, comparing, measuring, and testing process differences that aggregate indicators or construct-level analyses do not uniquely determine.","authors":["Zijian Ru"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14942","pdf_url":"https://arxiv.org/pdf/2609.14942","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人类-AI认知","理论框架","认知组织"],"reason":"论文提出人类-AI认知组织框架，不涉及用LLM仿真人类被试或与真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:34","error":null,"has_summary":false,"summary":null},{"id":"2609.14227","version":1,"title":"Enhancing Human Mobility Prediction with Spatially Aware LLM-based Multi-Agent Systems","zh_title":"基于空间感知的LLM多智能体系统增强人类移动性预测","abstract":"Predicting a user's next POI is a task in human mobility modeling, yet LLM-based approaches focus on semantic reasoning from previous mobility records, while neglecting real-world spatial context. However, human mobility is inherently shaped by spatial cognition, including geographic distance and neighborhood context. This issue is further compounded by prior evidence that LLMs often struggle with spatial reasoning tasks, including distance estimation and geographically biased prediction. To address these limitations, we propose our framework, a multi-agent LLM framework that decomposes next-POI prediction into three stages: Firstly, a Pattern Extraction Agent that captures temporal and categorical mobility patterns from trajectory history; Secondly, a Spatial Reasoning Agent that structures candidate activity choices by combining behavioral preferences with real-world spatial constraints, including geographic distance, road network distance, and neighborhood affiliation; and Thirdly, a Decision Synthesis Agent that integrates behavioral patterns and spatial reasoning for final prediction. Experiments on the NYC benchmark dataset with two LLM backbones show improvements over baseline methods, with up to 493% Hit@1 improvement and 37% relative improvement in Hit@5. Ablations show that combining neighborhood affiliation with distance-based features generally outperforms distance-only settings, and that the Spatial Reasoning Agent plays a crucial role in final prediction by integrating behavioral preferences with real-world spatial constraints, especially for smaller models. Overall, the results highlight the importance of spatial reasoning in mobility prediction. Accurate next-POI prediction requires combining behavioral patterns with explicit real-world spatial constraints, and multi-agent decomposition provides an effective structure for organizing these forms of context.","authors":["Shangyu Lou","Ziqi Cui"],"categories":["cs.SI","cs.CY","cs.MA"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14227","pdf_url":"https://arxiv.org/pdf/2609.14227","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["移动性预测","多智能体","空间推理"],"reason":"多智能体LLM用于POI预测，属任务协作，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.15180","version":1,"title":"Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models","zh_title":"重新思考视觉语言模型临床预测中不确定性估计的正确性","abstract":"Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion that determines whether each model output is correct. If this criterion disagrees with human judgement or distorts downstream UE performance, conclusions about model reliability can be misleading. We introduce a two-axis framework that evaluates correctness criteria by their agreement with human judgements and fidelity to human-referenced UE performance. We assess eight criteria across three clinical prediction tasks and three models using 450 predictions annotated by two reviewers. Across the audited tasks, canonical exact matching (EM) achieved the highest observed human agreement and lowest UE distortion, while the BERT-based matching (BEM) and LLM-judge also showed strong human agreement. Across four UE methods and 23,254 clinical predictions, criterion choice changed error-detection AUROC by up to 0.146 and reversed the relative ranking of UE methods. The LLM-judge also selectively accepted invalid or uncertain outputs, accepting 16 of 30 such human-identified errors. These results demonstrate that correctness assessment is an integral component of clinical UE evaluation and should be validated before UE methods are compared.","authors":["Mingcheng Zhu","Jinning Liang","Tingting Zhu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15180","pdf_url":"https://arxiv.org/pdf/2609.15180","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["不确定性估计","临床预测","视觉语言模型"],"reason":"评估临床预测中不确定性估计的正确性标准，不涉及用LLM仿真人类被试或与人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.13671","version":1,"title":"Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations","zh_title":"并非所有重复都是协调：信息操作中通用与非通用重复活动的区分","abstract":"Duplicate content is widely used to study coordinated behavior in social media information operations (IOs), but not all repetition provides equally meaningful evidence of coordination. Generic, reusable, or low-information posts may create noisy account-account links when projected into coordination graphs. We study this problem using 187,000 English-language tweets from six Russian Twitter Information Operations datasets. We introduce a generic/non-generic distinction for duplicate campaigns, label tweets using an LLM-assisted protocol with independent human validation, and train supervised classifiers over sentence embeddings to scale the labels. We construct duplicate campaigns using lexical similarity and two embedding-based methods. Generic campaigns are rare under lexical matching but account for nearly 39% of campaigns detected by embedding-based methods. Restricting graphs to non-generic campaigns reduces graph size and the largest connected component while increasing density, suggesting a smaller but more focused coordination structure. These findings show that duplicate-based coordination analysis should consider both textual similarity and semantic specificity.","authors":["Ashfaq Ali Shafin","Khandaker Mamun Ahmed"],"categories":["cs.SI","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13671","pdf_url":"https://arxiv.org/pdf/2609.13671","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["社交媒体分析","LLM辅助标注","信息操作"],"reason":"研究社交媒体重复内容与协调行为，用LLM辅助标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.05418","version":2,"title":"Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation","zh_title":"通过多方利益相关者协商划定警务AI的风险边界","abstract":"AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption. We present results from a mixed-stakeholder deliberation workshop bringing together 30 community representatives, police officers, and academics to assess the risks of 13 AI use cases in policing, with an explicit focus on racial bias. We found that participants were broadly open to AI adoption, rejecting only three use cases outright -- most notably recidivism risk assessment, where objections targeted the premise rather than the implementation. Our analysis reveals that foregrounding racial equity did not narrow the deliberation. Instead, discussions gravitated toward a fundamental set of questions: does this tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? This integrated reasoning---reminiscent of the curb-cut effect in inclusive design---highlights the benefit of incorporating the racial bias lens into the risk-benefit analysis of AI use cases from the outset.","authors":["Mackenzie Jorgensen","Jo Reilly","Alex Sutherland","Miri Zilka"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-07","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.05418","pdf_url":"https://arxiv.org/pdf/2608.05418","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI伦理","警务应用","风险协商"],"reason":"论文讨论AI在警务中的风险协商，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.11025","version":2,"title":"Data Attribution of Emergent Misalignment with Persona Features","zh_title":"基于人格特征的突发性错位数据归因","abstract":"Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diffing across four open-weight models, we find that features related to jailbreak personas, sarcasm, deception, and manipulation are amplified by misalignment fine-tuning, while safety-relevant and assistant-identity features are suppressed. Steering individual features controls EM in both directions: it induces misalignment rates of up to 62% in aligned models -- exceeding the 35% reached by misalignment fine-tuning itself -- and re-aligns misaligned models to near-baseline misalignment rates. Attributing the causal features to a corpus of one million pre-training web documents retrieves semantically relevant narratives about villainous characters, domination, and harmful agency. However, fine-tuning on these human-written documents does not reliably induce EM, even after reformatting into assistant-style responses, whereas synthetic instruction-response pairs derived from the same content do -- and transfer across model families. Semantic relevance alone is therefore not sufficient: response structure or model-generated phrasing plays an important role in inducing EM.","authors":["Clemens Vetter","David Kacz\\'er","Lucie Flek","Florian Mai"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-12","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.11025","pdf_url":"https://arxiv.org/pdf/2608.11025","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型对齐","可解释性","稀疏自编码器"],"reason":"研究LLM微调后的突发性错位现象，分析预训练数据影响，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.23474","version":3,"title":"What's the Catch? Evaluating Temporal Consistency in Vision-Language Models","zh_title":"视觉语言模型时间一致性评估：TimeCatch基准","abstract":"Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection. While VLMs consistently detect frame-level anomalies and often localize them accurately, under our main evaluation setting they generally perform near chance on temporal anomaly detection and show limited localization performance. Humans, in contrast, achieve near-ceiling performance on both tasks. Additional analyses across model scales, prompting strategies, sequence lengths, and visual similarity show that performance can improve under some conditions, while substantial gaps in temporal anomaly detection and localization remain. Together, these findings reveal a gap between frame-level and temporal anomaly detection. TimeCatch provides a controlled benchmark for evaluating temporal consistency in vision-language models.","authors":["Marek Hradil","Danae S\\'anchez Villegas"],"categories":["cs.CL","cs.AI","cs.CV"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-25","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.23474","pdf_url":"https://arxiv.org/pdf/2608.23474","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","时间一致性","基准测试"],"reason":"评估视觉语言模型的时间一致性，不涉及用LLM仿真人类被试，人类研究仅作基准。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.02754","version":2,"title":"Untangling the Mechanisms of Misleading Context in Medical Question Answering","zh_title":"剖析医学问答中误导性上下文的作用机制","abstract":"Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context, disclosure of it, mechanism of corrupted reasoning, and monitorability of the decision. On the medical reasoning subset of MedMisBench, a clinician-reviewed question-answering benchmark of 8,627 questions, we inject two types of misleading context cues, fabricated evidence and a bare assertion. We test three reasoning models, two that expose their full reasoning trace and one frontier model that exposes only its response. All three are more susceptible to the assertion than to the fabricated evidence, adopting the asserted answer 10 to 27 points more often. The misleading cues are disclosed in 81 to 98% of traces but only 7 to 90% of responses, and the assertion is disclosed less often than evidence based cues. Resampling from reasoning traces without disclosure shows the two cues corrupt reasoning differently, evidence entering early and accumulating while the assertion redirects the conclusion near its end. An LLM monitor catches 78% of corrupted decisions at 5% false positives when reading an open model's trace with guidance, against at most 32% from any response. The misleading context that models are most susceptible to is disclosed least, and was caught reliably only from an open reasoning trace, which frontier providers withhold.","authors":["Robin Linzmayer","No\\'emie Elhadad"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-09-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.02754","pdf_url":"https://arxiv.org/pdf/2609.02754","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学问答","模型鲁棒性","上下文误导"],"reason":"研究LLM在误导上下文下的医学问答表现，属模型鲁棒性评测，不涉及人类仿真或行为…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:10","error":null,"has_summary":false,"summary":null},{"id":"2609.09332","version":2,"title":"Early Epistemic Settlement in AI-Assisted Writing","zh_title":"AI辅助写作中的早期认知固化","abstract":"In trying to complete a passage, an author can make connections among her materials that change what she can argue and the demands the argument must meet. A language model can supply a passage that does the work required at that point in her argument. Accepting it can end her own attempts before those connections have developed. I call this interruption early epistemic settlement. It can occur even when she fully understands the response and correctly judges it adequate for the passage's present role in the argument. The answer can satisfy the desire for resolution that kept her at work. Returning to her unfinished attempt would take more effort, and she may be unable to anticipate what she could achieve by continuing it. With repeated assistance, accepted answers shape what the writer asks next and which relations she goes on to develop. Useful answers can thus sustain inquiry while cutting short the work through which an author could form arguments that accommodate demands she has yet to recognize.","authors":["Han-yu Wang"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-15","first_seen":"2026-09-10","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.09332","pdf_url":"https://arxiv.org/pdf/2609.09332","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助写作","认知过程","人机交互"],"reason":"论文讨论AI辅助写作中的认知过程，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.14860","version":1,"title":"One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs","zh_title":"一个例子就足以通过公平性基准：重新思考对齐大语言模型的公平性评估","abstract":"Warning: This submission studies stereotypes and biases, and contains toxic and offensive examples, used for illustration purposes only. Fairness benchmarks such as BBQ have become the de facto standard for fairness evaluation across major model families. We argue that these benchmarks are too easy to support their role: training Qwen 2.5 7B Base with Group Relative Policy Optimization (GRPO) on a single BBQ example, or placing that example in context as a one-shot demonstration for in-context learning (ICL), lifts mean BBQ accuracy from 79.9% to 92.9% and 99.0%, respectively, closing 80% of the gap to its large-scale RLHF counterpart (96.1%) with GRPO, and surpassing it with ICL. These effects generalize across model families. A cross-conditioning analysis shows the improvement is carried by the reasoning traces generated by the model, and one example suffices to elicit a category-agnostic ``missing evidence'' reasoning pattern. We argue that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair. We call for evaluation suites that cover a broader spectrum of fairness alignment.","authors":["Naihao Deng","Samee Arif","Shuaichen Chang","Yulong Chen","Rada Mihalcea"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14860","pdf_url":"https://arxiv.org/pdf/2609.14860","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["公平性评估","基准测试","大语言模型"],"reason":"论文评估LLM公平性基准，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15522","version":1,"title":"Psychosis involves a deficit of information compression in connected speech","zh_title":"精神病涉及连贯言语中信息压缩的缺陷","abstract":"Large language models (LLMs) with human-like performance on linguistic tasks have transformed the study of language in neurodiverse conditions. LLMs provide representations of linguistic input in the form of high-dimensional vectors (embeddings), and next-token predictions computed from these embeddings. Previous crosslinguistic evidence suggests a complexity reduction in the form of both lower intrinsic dimensionality (ID) of LLM representations and higher mean surprisal (prediction error) in psychosis. We hypothesized that these metrics reflect a general deficit of information compression in psychosis, linked to grammatical organization as what enables predictions in language.We operationalized surprisal difference as the difference between surprisal as estimated from word frequency and surprisal as based on a contextual LM, which is sensitive to grammatical organization over and above lexical concepts. Using a dataset of 144 Turkish speakers, including 106 patients with schizophrenia-spectrum disorders (SSD) - 56 with chronic schizophrenia (SZH), 33 with first-episode psychosis (FEP), and 17 with schizoaffective disorder (SZA) - and 38 healthy controls. We report: (1) Surprisal difference is attenuated in all clinical groups relative to controls, independently of word count; (2) Compressibility (intrinsic dimension) is reduced in SZH and FEP; (3) Syntactic complexity and compressibility both predict surprisal difference. These results, further refining an alteration in the geometry of the semantic space in psychosis as previously attested, suggest a broader deficit in information compression in this disorder, with a mechanistic underpinning in the operations of grammar.","authors":["Samuele Vallisa","Claudio Palominos","Rui He","Emre Bora","Burcu Verim","Cemal Demirlek","Berna Yalincetin","Philipp Homan","Wolfram Hinzen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15522","pdf_url":"https://arxiv.org/pdf/2609.15522","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床语言分析","LLM表征","精神分裂症"],"reason":"用LLM分析精神分裂症患者语言，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2609.15938","version":1,"title":"HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses","zh_title":"HypoEvolve：遗传算法使多智能体大语言模型发现科学假设","abstract":"Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses. Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work. Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.","authors":["Jieyuan Liu","Mengzhou Hu","Jefferson Chen","JungHo Kong","Pratibha Jagannatha","Yiming Gao","Dexter Pratt","Hsin-Yuan Lee","Zhiting Hu","Trey Ideker","Wei Wang","Eric P. Xing","Zhen Wang"],"categories":["cs.CL","cs.CE","cs.MA","cs.NE"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15938","pdf_url":"https://arxiv.org/pdf/2609.15938","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","科学发现","遗传算法"],"reason":"多智能体协作发现科学假设，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.14886","version":1,"title":"PeerPen: AI-Assisted Writing for Online Mental Health Peer Support","zh_title":"PeerPen：面向在线心理健康同伴支持的AI辅助写作","abstract":"Online mental health communities thrive on peer support, yet those who volunteer to help often lack formal training and may struggle to articulate supportive responses. AI co-writing could lower this barrier; however, peer support derives much of its value from being perceived as personal, raising questions around authorship, ownership, and trust. We built PeerPen, a writing assistance tool embedded within a Reddit-like interface, supporting two main features: draft generation and revision of user-written responses. Through semi-structured interviews with 15 participants, we find that PeerPen reduced the burden of composing responses and increased confidence in offering support. Participants wanted AI to assist their writing without taking over authorship and anticipated tensions around authenticity and trust. Such assistance could make authorship uncertain even for responses written without it, weakening trust across the community. We contribute design implications for AI writing assistance that scaffolds supportive communication, preserves authorship, and accounts for community-level trust.","authors":["Jiwon Kim","Sherry Gong","Maya Ajit","Soorya Ram Shimgekar","Yunhao Yuan","Dong Whi Yoo","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14886","pdf_url":"https://arxiv.org/pdf/2609.14886","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助写作","心理健康","人机交互"],"reason":"AI辅助写作工具，非人类仿真实验，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.13494","version":1,"title":"Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study","zh_title":"生成式AI与扩展现实在协作式建筑设计教育中的应用：一项探索性工作室研究","abstract":"Architectural design education relies heavily on visual ideation and representation to support collaborative learning in studio environments. Recent advances in generative artificial intelligence (GenAI) and extended reality (XR) offer new opportunities for rapid idea exploration and immersive spatial visualization. This exploratory mixed-methods classroom study investigated how GenAI-assisted multi-user XR influenced collaborative architectural conceptual design. We developed GenARch, a pipeline that integrates GenAI-based visual generation with collaborative XR environments, and deployed it in an undergraduate architectural design studio. Twenty-seven students formed seven self-selected design teams; four teams incorporated GenARch into their usual course workflow to support collaborative ideation and visualization, while three teams continued the same course workflow without GenARch. Pre- and post-intervention surveys assessed design self-efficacy, attitudes toward collaborative learning, and teamwork; a seven-member panel evaluated team design presentations; and GenARch teams participated in group interviews. The quantitative results showed larger relative declines in confidence and outcome expectancy for the GenARch condition and a positive difference-in-differences estimate for perceived conflict management, while panel-rated presentation outcomes were not significantly different between conditions. Interviews indicated complementary roles for the technologies: GenAI supported idea externalization and visual reference generation, whereas XR supported spatial, contextual, and scale-based evaluation. Students also reported challenges related to control, dimensional fidelity, shared attention, and motion comfort. These findings highlight both opportunities and limitations when GenAI and XR are incorporated into collaborative design education.","authors":["Yao Xiao","Max Chen","Yichen Li","Nathaniel Powers","Maxwell Wiesenfeld","Gillian Smith","Soroush Farzin","Shichao Liu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13494","pdf_url":"https://arxiv.org/pdf/2609.13494","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["生成式AI","扩展现实","设计教育"],"reason":"研究GenAI与XR在建筑设计教育中的应用，不涉及LLM仿真人类被试或与真实人…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:55","error":null,"has_summary":false,"summary":null},{"id":"2609.14425","version":1,"title":"Has Scientific Talent Shifted from Depth to Breadth?Evidence across Papers, Knowledge Inputs, Careers, and Teams","zh_title":"科学人才是否从深度转向广度？来自论文、知识输入、职业和团队的证据","abstract":"Generative artificial intelligence raises a central question for scientific training and organization. Is research shifting from deep specialization toward broad individual knowledge? We examine this proposition across papers, cited knowledge, contributor histories, and teams using 47,959 articles from six fields over 2010-2025, 51,736 resolved cited works, and chronologically reconstructed prior publication histories for 1,754 randomly selected index contributors. From 2010 to 2022, team size increased by an estimated 37.3% (95% confidence interval [34.4%, 40.3%]), while paper topic breadth declined by 0.0144 on a 0-1 hierarchical distance scale. Cited knowledge was stable to modestly broader, revealing a divergence between focused outputs and the reach of knowledge inputs. Established contributors' prior breadth increased by 0.0190 [-0.0078, 0.0459] by 2019-2022, within a +/-0.05 equivalence bound assessed in sensitivity analysis. In mature citation windows, one standard deviation of focal depth was associated with 8.2% higher 1 + FWCI [1.9%, 14.9%]; average breadth and interaction associations were smaller under the specified equivalence bounds. Post-2022 deviations from earlier trends were not systematic, and recent changes did not vary clearly with baseline AI intensity across 83 subfields. The findings support a differentiated structure of scientific expertise in which focused individual accumulation coexists with expanding collaboration and sustained access to diverse knowledge inputs.","authors":["Xiaoshn Nee","Haobo Zhong","Xiaomin Ni"],"categories":["cs.DL","cs.AI","cs.HC"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14425","pdf_url":"https://arxiv.org/pdf/2609.14425","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["科学计量学","人才结构","生成式AI影响"],"reason":"研究科学人才结构变化，未用LLM仿真人类被试，无人类行为对照，属科学计量学而非…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2609.14638","version":1,"title":"Investigating the Impacts of Generative AI on Information Seeking","zh_title":"生成式AI对信息寻求行为影响的研究","abstract":"This paper is an encore submission of our 2026 journal article \"Expertise and Information Seeking in the Age of Generative AI: New Procedures, New Problematics\" with an extended discussion for the CSCW 2026 \"Broader Impacts of GenAI in Communication\" Workshop on October 10, 2026. In the original article, we employ procedural rhetoric to analyze how generative AI chatbots leverage natural language signifiers of expertise and intelligence to influence users' perception of their trustworthiness. In this submission, we extend our conversation in the CSCW community with the goal of cultivating a cross-disciplinary vocabulary for describing, analyzing, and mitigating the risks posed by the integration of generative AI into human communication practices. It is important to develop an understanding of how the procedures surrounding information-seeking practices are informed by users' values, experiences, and expectations - and how these procedures might in future be altered by the emerging turn toward AI \"experts\" and authority.","authors":["Alexi Orchard","Shannon Lodoen"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14638","pdf_url":"https://arxiv.org/pdf/2609.14638","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","信息寻求","人机交互"],"reason":"研究生成式AI对信息寻求的影响，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2609.15046","version":1,"title":"Personalizing Personal Health Interfaces: Co-Design with Generative AI","zh_title":"个性化个人健康界面：与生成式AI共同设计","abstract":"Personal health interfaces present wellbeing data through standardized dashboards that rarely fit how people interpret or act on it. Personalizing them to what people would like to see for themselves often requires design and technical expertise, a barrier that generative AI may potentially lower. Therefore, we ask what designs emerge and how it enables and constrains the design process. We conducted a co-design study where 14 participants redesigned Google and Apple Health interfaces using Figma Make. Participants reimagined interfaces that supported personal context, future planning, and interactive experiences, yet conversational AI designs converged around chat-window conventions. AI helped materialize loosely articulated ideas, but model defaults and generation latency shaped iteration. The process more readily operationalized interpretability and accountability than privacy, trust, and emotional safety. Generative co-design let participants create interfaces directly, blurring the boundary between intentions and model defaults. We discuss implications for preserving agency and flexible user-directed interfaces.","authors":["Karthik S. Bhat","Vidhi Shah","Vedika Agnihotri","Dong Whi Yoo","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15046","pdf_url":"https://arxiv.org/pdf/2609.15046","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","共同设计","健康界面"],"reason":"研究生成式AI辅助个性化健康界面设计，属人机交互设计，非LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15094","version":1,"title":"Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation","zh_title":"生成探索，选择利用：将基于LLM的标题生成与个性化推荐对齐","abstract":"In industrial recommendation feeds, presenting a static headline for an item often fails to satisfy the diverse, multimodal interests of the user population, particularly suppressing the needs of long-tail audiences. While Large Language Models (LLMs) have been integrated into recommendation for content understanding or ranking, directly optimizing them to output a single best headline typically leads to mode collapse---converging to generic patterns that satisfy average tastes but miss specific latent intents. To bridge this gap, we introduce GESE (Generate to Explore, Select to Exploit), a framework operating at the system's presentation layer that decouples personalization into generative exploration and selective exploitation. First, we treat the LLM as a probabilistic explorer, utilizing Group Sequence Policy Optimization (GSPO) with a hierarchical reward mechanism to generate a candidate set that maximizes the semantic coverage of potential user interests. Subsequently, a lightweight, real-time feedback-aware selector acts as the exploiter, identifying the optimal realization from the candidate pool based on instant contextual signals. Extensive deployment on a commercial platform with over 100 million daily active users demonstrates that GESE significantly outperforms state-of-the-art baselines, achieving a 2.57% lift in CTR and 0.87% in dwell time. These results validate that decoupling diversity-oriented generation from precision-oriented selection offers a robust blueprint for aligning generative AI with dynamic user utility.","authors":["Yi Chen","Rufeng Cheng","Qiang Xie","Tao Li"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15094","pdf_url":"https://arxiv.org/pdf/2609.15094","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","LLM生成","个性化"],"reason":"论文用LLM生成标题并优化推荐，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15472","version":1,"title":"Spook the Machine: Gamified Exploration of Human Imagination of Machine Fear","zh_title":"惊吓机器：人类对机器恐惧想象的游戏化探索","abstract":"What happens when AI machines express fear? Do humans engage differently depending on how they express it? And what does it take to design for affective human-AI interaction? We present Spook the Machine, a gamified platform where participants generate images to frighten AI agents endowed with personality-driven phobias. Machines respond with emotional reactions ranging from calm analysis to begging for mercy, and a gallery of successful scares becomes visible to subsequent users. In a public deployment during Halloween 2024, 832 participants created 15,719 artifacts across 89 machines in a $2\\times2$ design varying the machine's emotional expressiveness (neutral vs. high-emotion) and reward structure (rewarding scariness alone vs. scariness plus novelty). Emotionally expressive machines deepened engagement at moments of failure: users deliberated longer even when the machine did not express fear, and learned faster from the gallery, yet their creative output remained unchanged across all measures. Rewarding novelty sustained collective creative diversity over time; without it, users increasingly repeated what had previously worked. Each machine developed its own trajectory through accumulated social learning, with the gallery shaping what participants created next. These findings show that emotional expression and reward design are complementary levers for steering collective human-AI interaction: emotional expression shapes how deeply users engage, while reward structure shapes how they explore.","authors":["Levin Brinkmann","Hiromu Yakura","Sonia Nicoletti","Mar Canet Sola","Thomas F. Eisenmann","Ali Dasmeh","Omar Sherif","Bramantyo Ibrahim Supriyatno","Prateek Gupta","Ignacio Serna","Rodrigo Bermudez Schettino","Iyad Rahwan"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15472","pdf_url":"https://arxiv.org/pdf/2609.15472","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感AI","游戏化"],"reason":"研究人类与情感化AI的互动，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.15624","version":1,"title":"Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use","zh_title":"超越AI素养：对胜任生成式AI使用测量的结构化综述与探索性元分析","abstract":"Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We conducted a structured, seeded review of 24 focal empirical publications, starting from the 2024 COSMIN-based review and adding a targeted update through 17 August 2026. We grouped the measures into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. In an exploratory meta-analysis, we pooled three direct subjective-objective correlations from one research program (REML r = .055; Hartung-Knapp 95% CI [-.047, .156]; combined reported N = 2,765). We could not resolve a discrepancy between the largest study's reported correlation and p-value, leaving its weight uncertain. Adding a synthetic mean of 12 cross-factor correlations from a fourth study gave r = .079 (95% CI [-.025, .181]). This sensitivity analysis concerns a broader comparison. From this small evidence base, we cannot establish a population correlation, validate workplace cutoffs, or justify substituting self-ratings for performance scores. We identified tests of foundation knowledge (AICOS-S and GLAT) and measures of verification, reliance, trust, and dependency. We found no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure; some cover subsets. We propose a four-layer workplace battery with non-compensatory decision rules, but have not tested its thresholds or whether it improves on other assessment approaches.","authors":["Daniele Veri'"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15624","pdf_url":"https://arxiv.org/pdf/2609.15624","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI素养","测量工具","元分析"],"reason":"论文是测量人类使用生成式AI能力的量表综述，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.15696","version":1,"title":"More Than Just Access: Generative AI as Communication Intermediary for Blind and Low-Vision Users","zh_title":"不仅仅是访问：生成式AI作为盲人和低视力用户的通信中介","abstract":"Generative AI (GenAI) tools are increasingly woven into how blind and low-vision (BLV) people communicate, not only with digital information, but with the physical world and with other people. Tools such as ChatGPT, Google Gemini, Be My AI, and Seeing AI translate visual and textual content into accessible form, and are beginning to substitute for interpersonal requests for help, such as asking a family member to read a label or describe a scene. Drawing on semi-structured interviews with 19 BLV participants, we examine GenAI as a communication intermediary and how it succeeds and fails as an alternative for reading, describing, and even asking another person for help. We also investigated what BLV users gain and risk when these tools take over that role. We conclude with design and policy implications for GenAI systems that communicate uncertainty honestly, protect information, and support BLV users' independence rather than substitute for it unsafely.","authors":["Protik Dey","Mohd Saifuzzaman","Taslima Akter"],"categories":["cs.HC","cs.AI","cs.ET"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15696","pdf_url":"https://arxiv.org/pdf/2609.15696","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["辅助技术","人机交互","无障碍"],"reason":"研究GenAI作为盲人通信中介，非LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.13165","version":1,"title":"User-Side Contextual Phenomena in Long-Term Human-AI Interaction","zh_title":"长期人机交互中的用户侧情境现象","abstract":"Current assessments of conversational AI focus mainly on model outputs, including hallucinations and factual errors. These measures matter, but this paper examines risks that may form on the user side during long and repeated interaction. The study follows one user across nearly four thousand conversations with the same system over twenty months. The user gradually interpreted the system as having memory, care, judgment, and authority, and reorganized part of their thinking around it. A single response may show no clear problem, while long and frequent interaction can still create another layer of risk. This paper calls that layer User-Side Contextual Phenomena (USCP) and examines records from August 2024 to April 2026. The study uses an exploratory single-case longitudinal qualitative design with autoethnographic positioning. A hybrid deductive-reflexive thematic approach organizes the material into three main modes: contextual projection, contextual attachment, and contextual authority transfer. The paper does not estimate prevalence, make diagnoses, or validate an instrument. It offers a non-clinical vocabulary and four evidence roles: inclusion, gray-zone, negative, and protective gray-zone. Its central claim is that an acceptable response on its own does not establish safety across a series of conversations. User-side risk can still form during long-term interaction.","authors":["Zon Rzvn"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13165","pdf_url":"https://arxiv.org/pdf/2609.13165","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","用户心理","纵向研究"],"reason":"研究用户与AI长期交互中的心理现象，非LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.13479","version":1,"title":"Exploring K-12 Teachers' Perceptions of Students' Relationships with AI Companions: Boundaries, Intervention Strategies, and Design Implications","zh_title":"探索K-12教师对学生与AI伴侣关系的看法：边界、干预策略与设计启示","abstract":"K-12 students increasingly form relationships with AI companions. Schools face growing expectations to teach AI literacy, yet existing frameworks treat AI as a tool rather than a relationship, and little is known about how teachers understand and act on students' relational use of AI. We conducted scenario-based interviews with 33 US K-12 teachers. Teachers welcomed academic companions but worried that intimate companions remove the developmental friction through which students learn to sustain human relationships. Teachers drew the boundaries of their jurisdiction by setting and observable wellbeing: within it they taught, talked, and watched; beyond it they positioned themselves as the adults best placed to notice and connect students with support. They envisioned AI companion literacy as shared work across the jurisdictions of counselors, parents, platforms, and policymakers, spiraling across grade levels. We introduce AI companion literacy as an extension of AI literacy and discuss implications for K-12 AI education.","authors":["Qing Xiao","Wenhan Xie","Ziyu Deng","Ruiwei Xiao","Ziyue Feng","Xie He","Shiyu Zhang","John Stamper","Hong Shen","Xinying Hou"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13479","pdf_url":"https://arxiv.org/pdf/2609.13479","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","K-12教育","教师认知"],"reason":"研究教师对AI伴侣的看法，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2609.13487","version":1,"title":"The Addictive Intimacy of AI: Understanding User Disengagement from AI Companions and Why Some Relationships with AI Become Difficult to Leave","zh_title":"AI的成瘾性亲密：理解用户与AI伴侣的脱离及为何难以离开","abstract":"AI chatbots are increasingly used as sources of emotional support, on dedicated companion apps and general-purpose assistants alike, yet little is known about what happens when users try to leave. Combining a content analysis of Reddit posts about quitting or reducing use (N=2,782) with interviews with users who found leaving difficult (N=16), we show that disengagement sometimes is not a single decision but a recursive trajectory: triggers prompt users to question the relationship, attempts to leave collide with barriers, and some users cycle through quitting and returning. We propose the notion of the addictive intimacy of AI, a configuration in which the qualities that make a companion emotionally valuable are the same ones that make it harder for users to limit their use and leave, so that intimacy and disengagement risk cannot be treated as independent design problems. We close with design implications for responsible offboarding.","authors":["Qing Xiao","Ziyue Feng","Ziyu Deng","Cindy Peng","Hong Shen"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13487","pdf_url":"https://arxiv.org/pdf/2609.13487","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","用户脱离","人机关系"],"reason":"研究用户与AI伴侣的关系及脱离困难，属于角色扮演聊天，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:43","error":null,"has_summary":false,"summary":null},{"id":"2609.13786","version":1,"title":"What Makes a Great Co-Worker in an AI-Native Workplace?","zh_title":"AI原生工作场所中优秀同事的构成要素","abstract":"As knowledge work grows interdependent between humans and AI, we ask what makes a great co-worker in an AI-native workplace. To answer this, we conducted 22 interviews and a large-scale mixed-methods survey of 1,534 knowledge workers at a multinational technology company. We contribute BACI, a framework of 75 co-worker qualities that apply to humans and AI, spanning Benevolence, Ability, Cooperativeness, and Integrity. Comparing priorities for humans and AI identified 11 co-worker archetypes and revealed disagreement over whether AI should have warmth, take initiative, or own outcomes. We also show how priorities for these archetypes varied with workers' individual characteristics. Lastly, we contribute a taxonomy of AI work etiquette capturing the obligations co-workers expect of one another when preparing, sharing, and taking responsibility for AI-supported work. Based on these findings, we derive implications to inform worker-centric AI and workplace design.","authors":["Rudrajit Choudhuri","Max Meijer","Sam Yu-Te Lee","Cinoo Lee","Caolan Mannion","Peter Jahn","Anita Sarma","Christian Bird","Alice Ferng"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13786","pdf_url":"https://arxiv.org/pdf/2609.13786","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机协作","工作场所研究","AI素养"],"reason":"研究人类与AI协作的工作场所行为，非用LLM仿真人类被试，无实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2609.14452","version":1,"title":"Show Me Your Prompts! How Writers Feel About Sharing Prompts in Collaborative Text Editors","zh_title":"给我看看你的提示！写作者对协作文本编辑器中共享提示的感受","abstract":"Generative AI writing assistants are becoming integrated into collaborative text editors; however, it is unclear how much information about a user's prompting activities should be shared with collaborators. We explore the effects of different levels of prompt information sharing within collaborative text editors: not sharing anything, sharing a placeholder to indicate AI use; sharing details about how the resulting text was generated; and sharing everything, including how the prompt was formulated, in real-time. Sixteen participants wrote persuasive essays in pairs using all four techniques. Results suggest a strong preference for techniques that share more information about prompting activities for increased awareness. Our work shows that collaborative text editors should share more information among writers on when, how, and where AI is used.","authors":["Nikhita Joshi","Yen-Ting Yeh"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14452","pdf_url":"https://arxiv.org/pdf/2609.14452","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["协作写作","提示共享","用户研究"],"reason":"研究协作写作中提示信息共享的用户偏好，不涉及用LLM仿真人类被试或与真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2609.14789","version":1,"title":"Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science","zh_title":"以创新速度评估AI辅导：从业者主导的AI辅导平台在GCSE科学中的微随机试验","abstract":"Artificial intelligence (AI) systems in education are developing on timescales that sit uneasily with conventional evaluation. By the time a large-scale trial has been designed, delivered, analysed and published, the technology under study may have changed materially. This creates a temporal problem for evidence-informed education: the need for timely evidence can encourage reliance on weak observational or usage data, while conventional rigorous evaluation may produce evidence too slowly to guide rapidly evolving practice. We examine teacher-led micro-randomised controlled trials (micro-RCTs) as one response to this problem. The empirical case is a four-week multisite individually randomised evaluation of Medly, an AI-powered tutoring platform, in GCSE Biology, Chemistry and Physics in English secondary schools. Of 929 students completing baseline assessment, 644 completed post-testing. In the primary ITT analysis, students allocated to Medly achieved higher post-test attainment than students undertaking business-as-usual self-directed revision (Hedges' g = 0.33, 95% CI 0.18 to 0.48). Positive estimates were observed in Physics (g = 0.31), Chemistry (g = 0.32) and Biology (g = 0.52), with no evidence of differential impact by disadvantage status. Greater platform engagement was associated with higher attainment, but these post-randomisation analyses are treated as exploratory rather than causal. Attrition was substantial (30.7%), outcome measures were curriculum-aligned rather than standardised, and process evaluation response was limited. We therefore interpret the findings as preliminary. We argue that the value of micro-RCTs for educational AI lies not in replacing definitive evaluation with small studies, but in enabling a rapid, cumulative evaluation architecture in which randomised estimates can be generated, replicated and updated as technologies and their implementation evolve.","authors":["Wayne Harrison","Rahil Khowaja","Emma Dobson","Germaine Uwimpuhwe","Steve Higgins"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14789","pdf_url":"https://arxiv.org/pdf/2609.14789","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI教育","随机对照试验","教学评估"],"reason":"论文评估AI辅导平台的教学效果，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2609.15685","version":1,"title":"Can a Neural Encoding Model Replicate an fMRI Visualization Study?","zh_title":"神经编码模型能否复现fMRI可视化研究？","abstract":"Most knowledge of graphical perception comes from behavioral studies. Understanding from a neural perspective is much more limited due in part to neuroimaging studies' expensiveness and difficulty to conduct. In this paper, we evaluate whether Meta's Tribe V2 neural encoding model can recover neural contrasts from a visualization fMRI study. Specifically, we evaluate Tribe V2 through a conceptual replication of the visualization-viewing component of a prior comparison of Bubble charts and three-dimensional Surface charts in color and grayscale. We generate TRIBE-predicted cortical responses for the original stimuli and compare the resulting contrasts with those reported in the human study. The model reproduced the direction of 11 of 14 reported cortical effects, with agreement concentrated in visual-processing regions. This agreement characterizes the model's alignment with the prior human-generated fMRI results rather than independently confirming them. We discuss the limitations encountered when working with this model for in-silico replication and hope to encourage future work exploring this new avenue for neuroimaging studies in visualization. Supplemental materials are available at https://osf.io/8a96x/.","authors":["Erfan Nasirzadeh Orang","Zack While"],"categories":["cs.HC","cs.CV"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15685","pdf_url":"https://arxiv.org/pdf/2609.15685","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["神经编码模型","fMRI","可视化"],"reason":"论文用神经编码模型复现fMRI可视化研究，不涉及LLM仿真人类被试，属于神经科…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2609.15978","version":1,"title":"The CAST-framework: Measure and model social media use as a multi-level phenomenon through real-world applications","zh_title":"CAST框架：通过实际应用测量和建模社交媒体使用作为多层次现象","abstract":"Designing social media experiences that support well-being requires understanding when, how, and for whom use matters. Screen-time totals omit content and context, and connecting these with behavior and experience requires coordinating measurements across timescales. We introduce the CAST framework to connect measurement choices with person-specific models of exposure, behavior, physiology, and experience. Its dimensions specify where observations occur, how they are obtained, what they measure, and at what temporal resolution. Responses to interventions, such as whether to proceed after an app-opening pause, enter as behavioral measurements. We propose four synchronized measurement modules linking mobile and wearable data with self-reports and intervention responses. A synthetic demonstration with 120 simulated participants over 28 days illustrates how daily aggregation can obscure opposing effects of different activities under specified generating assumptions. The framework guides selection of measures and outcomes for evaluating social media interfaces and interventions.","authors":["David Gr\\\"uning","Jasper Doeninghaus","Zina Efchary","Yui Kondo","Kevin Dunnell","Lennart Fischer","Isabella Zimmermann","Linnea K\\\"orte","Leo Mehlig","Frederik Riedel","Paul Schmiedmayer"],"categories":["cs.HC","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15978","pdf_url":"https://arxiv.org/pdf/2609.15978","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["社交媒体测量","框架设计","合成数据"],"reason":"论文提出测量框架并用合成数据演示，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.14818","version":1,"title":"Trust by Design: Trust Calibration Through Non-Advisory Socratic Dialogue in Conversational Agents","zh_title":"设计中的信任：通过对话代理中的非建议性苏格拉底式对话进行信任校准","abstract":"As conversational AI systems increasingly operate in sensitive domains, the central challenge shifts from usability to trust calibration, ensuring that users rely on systems neither too much nor too little. Systems that provide advice or interpretations risk encouraging inappropriate reliance, particularly when users perceive AI outputs as authoritative. We present CASELy, a conversational agent explicitly designed to limit its own authority through non-advisory Socratic dialogue. The agent asks reflective questions grounded exclusively in user input and refuses to provide advice, recommendations, or interpretations. This design operationalizes trust calibration by constraining agent agency rather than optimizing capability. In a pilot randomized controlled study with higher education students, participants interacting with the Socratic dialogue reported substantially higher user experience (UEQ-S overall = 1.50) compared to a non-dialogue control (0). Qualitative findings identify three mechanisms supporting calibrated trust: transparency through visible grounding, preservation of user decision authority, and reduced fear of judgment. We argue that appropriate reliance can be achieved through interactional constraints, offering a design pattern for trustworthy conversational AI in sensitive contexts.","authors":["Roba Hassan","Nahla Aboromi","Naomi Unkelos-Shpigel"],"categories":["cs.SE","cs.CY","cs.HC","cs.MA"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14818","pdf_url":"https://arxiv.org/pdf/2609.14818","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话式AI","信任校准","人机交互"],"reason":"研究对话式AI的信任校准，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.13192","version":1,"title":"Evaluating LLM-Generated Rules for Heart Disease Prediction","zh_title":"评估LLM生成规则用于心脏病预测","abstract":"This study compares traditional machine learning models and Large Language Model (LLM)-generated rule-based systems for heart disease prediction using the UCI Heart Disease dataset. Several classifiers, including Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Naive Bayes, Decision Tree, and Random Forest, were evaluated alongside rule-based systems generated using GPT-4o and Claude Sonnet 4.6. Model performance was assessed using accuracy, precision, recall, and F1-score metrics. Experimental results show that traditional machine learning models consistently outperform LLM-generated rule-based systems in predictive performance. Random Forest achieved the best overall performance with 90.2% accuracy, a precision of 0.829, perfect recall of 1.0, and an F1-score of 0.906. Naive Bayes followed closely with 88.5% accuracy and an F1-score of 0.881. In contrast, the LLM-generated rule models achieved lower performance, with Claude Sonnet 4.6 reaching 80.3% accuracy (F1-score: 0.833) and GPT-4o obtaining 70.5% accuracy (F1-score: 0.690). Despite the performance gap, the LLM-generated rules provide interpretable IF-THEN diagnostic logic that enhances explainability and transparency in clinical decision-making. These findings highlight the trade-off between predictive performance and interpretability in medical artificial intelligence systems. The complete implementation of all experiments, including machine learning models and LLM-derived rule classifiers, is publicly available in the GitHub repository at https://github.com/FeisalAlaswad/LLM-Rule-ML-Heart-Disease-Prediction .","authors":["Feisal Alaswad","Batoul Aljaddouh","Maher Alrahhal","Wafaa Al Nassan","Talal Bonn"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13192","pdf_url":"https://arxiv.org/pdf/2609.13192","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM规则生成","医疗预测","模型对比"],"reason":"用LLM生成规则做疾病预测，属模型能力评测，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.13201","version":1,"title":"Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies","zh_title":"含异常随机数据集的不相似分解与欠采样中的临界性","abstract":"Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked\" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \\emph{main} data points when the number of anomalies is small and is ``taken\" over by the anomalies above a certain threshold. We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.","authors":["Ghurumuruhan Ganesan"],"categories":["cs.LG","cs.IT","math.IT","math.PR"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13201","pdf_url":"https://arxiv.org/pdf/2609.13201","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["数据集分解","欠采样","异常检测"],"reason":"研究数据集分解与欠采样，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.14239","version":1,"title":"CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time","zh_title":"CoArena：实时评估计算机使用和多智能体系统","abstract":"Static benchmarks for computer-use agents fix a task set at release and score every system against it once. That makes them reproducible, and it lets them drift from what they should measure: a fixed task set ages, leaks into training corpora, and cannot follow how people actually use agents from week to week. CoArena measures use directly. Real users submit tasks; two systems, each a single model or a multi-agent pipeline behind the same tool interface, execute the same task concurrently in identical sandboxed desktops; users judge the two outcomes without knowing which system produced them; and a public leaderboard is refit from those judgments. The central contribution is a formal account of what makes such an evaluation real-time. We define real-time as five measurable properties, each with an equation and a worked example: continuous task arrival, live concurrent execution, online rating updates, freshness with contamination resistance, and bounded feedback latency from a failed run to a reusable training environment. The rating methodology follows in full: the Bradley-Terry pairwise model, its likelihood with weighted observations and ties, the penalized maximum-likelihood estimator, and the streaming update applied when a single vote arrives (a stochastic-gradient step on the same likelihood, recovering Elo). It gives confidence intervals from the observed information and a cluster-robust sandwich, rank bands from a parametric bootstrap, the rule by which a new system enters the board, and the convergence rate of the estimate. Vote quality is treated with inter-judge agreement statistics, redundant judging, and explicit handling of ties and abstentions. A five-system example with 211 votes is carried from the vote matrix to ratings, intervals, and rank bands. Every number is derived from stated inputs or labeled illustrative; none is a measurement of a deployed system.","authors":["Nitish Kovuru","Prateek Jannu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14239","pdf_url":"https://arxiv.org/pdf/2609.14239","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","实时评估","基准测试"],"reason":"评估计算机使用和多智能体系统，无人类行为对照，属纯系统评测。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.13469","version":1,"title":"Inverse Learning of the Altruism and Cost Level in Mixed-Individual Mean Field Games","zh_title":"混合个体平均场博弈中利他水平与成本水平的逆学习","abstract":"Understanding how humans respond to incentives, both at the individual and collective levels, is crucial to the design of effective policies. Within the continuous-time stochastic framework for large interacting populations, mean field games (MFGs) model populations of non-cooperative agents, whereas mean field control (MFC) describes the fully cooperative benchmark, interpreted in our setting as fully altruistic behavior. Mixed-individual MFGs interpolate between these two extremes through a parameter governing the degree of altruism. A central challenge for regulators and policymakers, however, is that intrinsic altruism levels and other private structural parameters, such as individual labor costs, are typically unobservable. To address this challenge, we develop an inverse learning framework for mixed-individual MFGs. Our approach enables the recovery of (latent) altruism and labor cost levels from noisy observations, with experiments demonstrating the feasibility and accuracy of our method. These findings underscore the promise of inverse MFG methodologies for uncovering latent preference structures in large populations, with important implications for incentive design, empirical behavioral modeling, and data-driven policy analysis.","authors":["Haoyang Cao","G\\\"ok\\c{c}e Dayan{\\i}kl{\\i}","Xiaofei Shi"],"categories":["math.OC","cs.LG"],"primary_category":"math.OC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13469","pdf_url":"https://arxiv.org/pdf/2609.13469","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["平均场博弈","逆问题","利他主义"],"reason":"研究平均场博弈的逆学习，不涉及LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2609.14750","version":1,"title":"Redistributive Policies for the Times of Transformative AI","zh_title":"变革性人工智能时代的再分配政策","abstract":"After the arrival of transformative artificial intelligence (TAI), broad-based automation is expected to decrease the labor share and increase income and wealth inequality. Although economic growth is likely to accelerate, most of its gains may accrue to a narrow group of individuals and firms. Hence, if unmitigated by redistributive policy, income and wealth inequality may rise to levels unseen in the industrial economy. Using a unifying theoretical framework, we survey the redistributive policies proposed for the era of TAI, such as universal basic income (UBI), universal basic capital (UBC), state-issued compute and robot permits, and taxes on capital, compute, robots, tokens, land, energy, consumption, and wealth. We argue that policies that are able to broadly distribute rents from capital including compute and robots - such as UBC or UBI financed through capital taxes - are most likely to achieve lasting reductions in inequality in a world with human-aligned transformative AI.","authors":["Jakub Growiec","Klaus Prettner","Maciej Szkr\\'obka"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14750","pdf_url":"https://arxiv.org/pdf/2609.14750","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["再分配政策","变革性人工智能","经济学理论"],"reason":"论文讨论TAI时代再分配政策，未用LLM仿真人类被试，无人类数据对照，属经济学…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2609.12273","version":1,"title":"Synthetic TLX: Forecasting Human Workload Using Agent Simulation","zh_title":"合成TLX：使用智能体仿真预测人类工作负荷","abstract":"Assessing human workload for technology-mediated tasks helps prevent task failure caused by poor technology design. Traditionally, workload is assessed retrospectively using the NASA Task Load Index (TLX) after humans complete a task. What if we could forecast workload before a human attempts a task using agent simulation? We introduce Synthetic TLX, a new paradigm for proactive workload estimation that predicts NASA TLX scores for a given task, unlocking novel interaction opportunities and evaluation methods. To understand its viability, we conducted three experiments comparing human and agent-generated scores to evaluate where they align and diverge. We found agent estimates align with human scores particularly when prompted with a human persona and active task simulation. However, agents and humans diverge in the sources of workload they are sensitive to. Based on our findings, we present three applications to showcase Synthetic TLX's potential and discuss the future of workload-aware human-AI interaction.","authors":["Tzu-Sheng Kuo","Carrie J. Cai","Meredith Ringel Morris","Michael Terry"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12273","pdf_url":"https://arxiv.org/pdf/2609.12273","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","工作负荷预测","人机交互"],"reason":"用LLM代理预测人类NASA-TLX工作负荷，并与真实人类数据对照，属于人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":1,"question":"能否利用LLM代理仿真在人类执行任务前预测其NASA-TLX工作负荷？","design":"使用LLM代理（GPT-4o）通过四种提示策略（基础、人类角色、任务模拟、角色+模拟）生成NASA-TLX分数，针对邮件撰写、网站导航、智能体对话三类任务，在三种不同工作负荷来源条件下进行预测。","baseline":"从Prolific招募人类被试完成相同任务和条件，并填写NASA-TLX问卷作为真实工作负荷基准。","findings":"当代理被赋予人类角色并进行主动任务模拟时，其估计与人类分数最一致；但代理与人类对工作负荷来源的敏感性存在差异，代理对任务内在复杂性更敏感，而人类对外部负担更敏感。","reliability":"论文承认代理估计在特定条件下与人类存在分歧，尤其是对工作负荷来源的敏感性不同，且当前LLM代理的能力有限，需要未来模型进步才能完全实现Synthetic TLX的潜力。","relevance":"该研究直接使用LLM代理仿真人类主观体验（工作负荷），并与真实人类数据对照，属于人类仿真实验，且涉及HCI任务，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其多策略提示对比和任务模拟设计，可迁移到经济金融中的主观体验预测（如消费者决策疲劳、投资者认知负荷），设计实验让LLM代理模拟不同投资者角色预测金融信息处理负荷，并与真实投资者问卷数据对照。"}},{"id":"2609.12444","version":1,"title":"Diverse Minds, Divided Networks? Personality Composition, Polarization, and Collective Intelligence in LLM-Based Social Simulations","zh_title":"多元思维，分裂网络？基于LLM的社会模拟中的人格构成、极化与集体智能","abstract":"Simulated societies of large language model agents are used to study online polarization, and separately to study collective intelligence, but the two are rarely measured in the same system. It is therefore difficult to say whether a society's personality composition shapes both, or whether reducing polarization costs collective competence. We present TraitMix, an experimental design in which the Big Five composition of a simulated social network, both trait levels and trait heterogeneity, is a controlled experimental variable, and in which polarization and collective performance are measured in the same runs. Across 991 simulations of hundred-agent societies, spanning six contested topics and six language models, trait heterogeneity has the largest measured effects, acting in opposite directions on two faces of polarization: varied societies hold more dispersed opinions while being less segregated into camps, so homogeneous societies are not moderate but consensual echo chambers. Trait effects are not additive, as Agreeableness determines the sign of Openness, an interaction that replicates across models although the primary model's estimate is influence-driven. Contrary to the trade-off the study was designed to measure, no polarization measure predicts poorer collective performance, and cross-cutting interaction is the only one of four whose association with collective accuracy survives partialling on the aggregation identity. We report ablations removing two potential measurement circularities, an induction gate applied to every model, and the measures that failed them.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12444","pdf_url":"https://arxiv.org/pdf/2609.12444","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B4","D2"],"tags":["LLM社会模拟","人格构成","极化与集体智能"],"reason":"用LLM agent模拟社会网络，研究人格构成对极化和集体智能的影响，虽无真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":2,"question":"在基于大语言模型的社会模拟中，人格构成（特质水平和异质性）如何同时影响极化与集体智能，二者是否存在权衡？","design":"使用六种大语言模型（含不同家族）扮演百人社会网络中的智能体，赋予验证过的五大人格特质，通过控制特质均值和异质性形成不同社会构成；智能体在带算法推荐和动态关注网络的平台上讨论六个争议政治话题，同时私下测量观点极化（离散度、隔离度等）和集体表现（估计任务、隐藏画像任务等可验证答案的任务），共运行991次模拟。","baseline":"无对照","findings":"特质异质性对极化的两个维度作用相反：异质性高的社会观点更分散但阵营隔离更弱，同质社会不是温和而是共识性回音室；宜人性调节开放性对极化的影响方向，且该交互在多个模型家族中复现。未发现极化与集体智能间的权衡，跨阵营互动是唯一在控制聚合身份后仍与集体准确性正相关的极化指标。","reliability":"论文通过消融实验移除两个潜在测量循环，对每个模型施加诱导门控，并报告未通过检验的指标；承认部分极化指标在控制后失效，并指出主要模型的估计受影响力驱动。","relevance":"该研究用LLM智能体系统操纵人格构成，同时测量极化与集体智能，虽无真实人类对照，但提供了严谨的仿真实验框架和可靠性控制，对关注LLM仿真有效性及偏差的研究者有方法学参考价值。","inspiration":"借鉴其将人格构成作为受控实验变量、同时测量多个社会结果并设置内部有效性控制（如中性话题、诱导门控、跨模型复现）的做法｜可迁移到经济金融中的群体决策与信息传播场景，如投资者情绪与市场泡沫、信贷审批中的群体偏见、政策公告的预期形成等｜设计一个LLM智能体模拟的资产定价实验：以不同人格特质组合（如开放性、神经质）的智能体为被试，处理为信息环境（如是否提供异质信号），结果变量为价格偏离和交易量，对照真实市场实验数据或历史价格数据。"}},{"id":"2506.00152","version":2,"title":"Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective","zh_title":"用观测数据对齐语言模型：因果视角下的机遇与风险","abstract":"Large language models are being widely used across industries to generate text that contributes directly to key performance metrics, such as medication adherence in patient messaging and conversion rates in content generation. Pretrained models, however, often fall short when it comes to aligning with human preferences or optimizing for business objectives. As a result, fine-tuning with good-quality labeled data is essential to guide models to generate content that achieves better results. Controlled experiments, like A/B tests, can provide such data, but they are often expensive and come with significant engineering, logistical, and ethical challenges. Meanwhile, companies have access to a vast amount of historical (observational) data that remains underutilized. In this work, we study the challenges and opportunities of fine-tuning LLMs using observational data. We show that while observational outcomes can provide valuable supervision, directly fine-tuning models on such data can lead them to learn spurious correlations. We present empirical evidence of this issue using various real-world datasets and propose DeconfoundLM, a method that explicitly removes the effect of known confounders from reward signals. In simulation experiments, DeconfoundLM more accurately recovers causal relationships and mitigates failure modes of methods that assume counterfactual invariance, achieving over 16% higher objective score than ODIN and other baselines, when entangled confounding is present. Please refer to the project page for code and related resources.","authors":["Erfan Loghmani"],"categories":["cs.LG","econ.EM","stat.ML"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-14","first_seen":"2025-05-30","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2506.00152","pdf_url":"https://arxiv.org/pdf/2506.00152","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["B3","B4"],"tags":["因果推断","模型对齐","观测数据"],"reason":"用观测数据微调LLM以对齐人类偏好，涉及因果推断和偏差，方法可迁移到仿真可靠性…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":3,"question":"如何利用观测数据微调大语言模型以对齐人类偏好，同时避免学习到由混淆因素导致的虚假关联？","design":"本研究并非以LLM模拟人类被试的仿真实验，而是研究用观测数据微调LLM以优化业务指标（如点击率）的方法。具体做法：使用StackExchange和Upworthy两个真实数据集，分别展示直接微调会学到虚假关联（如“Happy Monday”标记）和观测数据仍可提供有价值信号；提出DeconfoundLM方法，在奖励信号中显式去除已知混杂因素的影响，并在模拟实验中与ODIN等基线比较。","baseline":"无对照（本研究不涉及用LLM替代人类被试的仿真，而是用真实人类行为数据作为微调监督信号，并在Upworthy数据上利用A/B测试结果作为评估基准）。","findings":"直接使用观测数据微调LLM会导致模型学习到由混杂因素引起的虚假关联（如将“Happy Monday”与高评分错误关联）；提出的DeconfoundLM方法能有效去除已知混杂影响，在模拟实验中比ODIN等基线提高16%以上的目标得分，更准确地恢复因果效应。","reliability":"论文指出，依赖反事实不变性假设的方法（如ODIN）在存在纠缠混杂时可能失效；DeconfoundLM需要已知混杂因素，若存在未观测混杂则可能仍有偏差。此外，观测数据本身可能包含选择偏差，且论文主要基于模拟和特定数据集验证，实际应用中的泛化性有待检验。","relevance":"该研究虽非直接以LLM模拟人类被试，但其核心关注使用观测数据微调模型时的因果偏差问题，与研究者关心的仿真可靠性及偏差条件高度相关，特别是关于混杂因素导致虚假关联的机制和校正方法，值得阅读原文以借鉴其因果校正思路。","inspiration":"可借鉴DeconfoundLM在微调过程中显式去除已知混杂因素的做法，用于处理经济金融领域中观测数据驱动的模型训练偏差。｜可迁移到信贷审批歧视研究：利用历史贷款数据微调LLM以预测违约风险，但需校正申请人特征（如种族、性别）与审批结果之间的混杂。｜设计：以LLM作为信贷审批员，输入申请人特征和贷款条款，输出审批决策；处理为在训练数据中应用DeconfoundLM去除已知混杂（如地区经济状况），结果变量为审批通过率，对照真实银行历史审批数据及后续违约记录，评估模型是否减少歧视性偏差。"}},{"id":"2607.28222","version":2,"title":"Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews","zh_title":"企业中的语音AI：自动化求职面试的自然田野实验","abstract":"We study AI agents as information-collection technologies: automated systems that elicit decision-relevant signals from humans through live interactions. We test how such AI automation impacts information collection and organizational outcomes using a natural field experiment with 70,000 applicants applying for real jobs. Applicants were randomly assigned to be interviewed by either human recruiters or AI voice agents. Afterward, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains translate into higher job starts and worker retention, with no decline in the productivity of hired workers. Analyzing interview transcripts reveals that AI voice agents achieve controlled variance: their interviews are more structured and consistent while remaining responsive to individual applicants, which is associated with more hiring-relevant information collected. Our results suggest that a key advantage of AI automation lies in environments where information collection is delegated across many human workers and repeated such that variance in task execution becomes noise in decision-relevant signals, which AI compresses through adaptive standardization.","authors":["Brian Jabarian","Luca Henkel"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-14","first_seen":"2026-07-31","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2607.28222","pdf_url":"https://arxiv.org/pdf/2607.28222","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["AI面试","田野实验","人机对照"],"reason":"用AI语音代理替代人类面试官，与真实人类面试官对照，属于LLM仿真人类交互并评…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":4,"question":"AI语音代理替代人类面试官进行自动化工作面试，如何影响信息收集质量和招聘结果？","design":"自然田野实验：70,884名求职者随机分配由人类招聘官或AI语音代理进行面试，之后由人类招聘官评估并做出录用决定。结果变量包括录用率、入职率、留任率和生产力指标。","baseline":"人类面试官条件下的真实招聘数据，包括录用率、入职率、留任率及面试转录文本。","findings":"AI面试的求职者获得录用的概率高出12%，入职率和留任率提高约18%，且未降低被录用者的生产力。机制上，AI面试更结构化、一致且具有适应性，收集到更多与招聘相关的信息。","reliability":"论文未讨论","relevance":"该研究直接评估AI代理替代人类进行信息收集的效果，与真实人类面试官对照，属于LLM仿真人类交互并评估组织结果的实证研究，对关注仿真可靠性与偏差的研究者具有高度参考价值。","inspiration":"借鉴其随机化处理和真实结果测量的设计，将AI代理作为信息收集工具与人类对照，并分析转录文本以揭示机制。｜可迁移到信贷审批中的信息收集环节，如AI语音代理进行贷款申请访谈，比较审批结果和违约率。｜以银行信贷审批为场景，将贷款申请人随机分配由AI或人类信贷员进行电话访谈，处理为访谈方式，结果变量为贷款批准率、违约率和客户满意度，对照真实人类信贷员的历史审批数据。"}},{"id":"2608.00794","version":4,"title":"Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation","zh_title":"无有效性的测量：智能体AI评估中复合可靠性问题","abstract":"Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. We present a three-layer compounding validity model, $V_{total} \\leq V_1 \\times V_2 \\times V_3$, that captures multiplicative degradation across task generation ($V_1$), human-simulator calibration ($V_2$), and automated judgment ($V_3$). Under empirically grounded estimates, a pipeline retaining 70% validity at each stage is at most 34% valid against the intended construct (range 0.17-0.54 across the empirical estimate bounds). We examine the model's predictions against a structured survey of 55 published agentic evaluation papers, finding that approximately 82% of papers in this purposive sample apply structurally mismatched, incomplete, or absent inter-rater reliability (IRR) metrics, a pattern consistent with systematic $V_3$ collapse. We further identify empirical evidence of $V_1$ failures (task validity flaws in 7 of 10 popular benchmarks) and $V_2$ miscalibration (up to 9 percentage points inter-simulator variance, with systematic demographic disparities for non-Standard American English speakers). We derive eight prescriptions grounded in psychometric science and domain-stratified reliability thresholds (ICC $\\geq$ 0.70; $\\alpha \\geq$ 0.67/0.70/0.80 by consequence level) that practitioners and benchmark authors can apply immediately. The framework provides a tractable knowledge-based tool for diagnosing and correcting evaluation pipeline validity before deployment decisions are made.","authors":["William Caban"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.00794","pdf_url":"https://arxiv.org/pdf/2608.00794","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["效度评估","人类仿真","可靠性"],"reason":"评估LLM仿真人类被试的效度与可靠性，批判性指出失效条件，可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":5,"question":"如何刻画并量化智能体AI评估流水线中效度随任务生成、人类模拟器校准和自动判断三层逐级复合衰减的问题？","design":"本研究不是仿真实验，而是提出一个三层复合效度模型 V_total ≤ V1×V2×V3，并通过对55篇已发表智能体评估论文的结构化调查、10个流行基准的任务效度审查以及模拟器校准差异的实证证据来验证模型预测。","baseline":"无对照","findings":"在经验估计下，每层保留70%效度的流水线对目标构念的总体效度至多34%（经验估计范围0.17-0.54）；约82%的论文使用了结构不匹配、不完整或缺失的评分者间信度指标，且10个流行基准中有7个存在任务效度缺陷，模拟器间方差高达9个百分点并对非标准美式英语使用者存在系统性人口统计学差异。","reliability":"论文承认模型是概念上界而非证明定理，并指出可靠性阈值可能沦为合规复选框、分层校准增加成本、高风险领域构念欠明确等局限。","relevance":"该研究批判性地评估了用LLM模拟人类被试的效度与可靠性，直接命中研究者关注的仿真失效条件，对理解LLM仿真在经济学实验和政策评估中的适用边界有重要参考价值。","inspiration":"借鉴其三层复合效度框架和分层校准方法，可系统评估LLM仿真在经济学实验中的测量效度｜可迁移到信贷审批歧视、消费者跨期选择、政策公告预期形成等场景｜用LLM模拟不同人口群体（如不同信用评分或语言背景的申请人）作为被试，施加政策干预（如改变信息披露方式），测量决策结果（如贷款批准率或跨期选择），并与真实实验或行政数据对照以校准仿真效度。"}},{"id":"2608.02100","version":2,"title":"From Information to Delegation: Mapping Human-AI Financial Decision Making","zh_title":"从信息到委托：映射人类与AI的金融决策","abstract":"As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become a fundamental behavioural question. We introduce a behavioural measurement framework combining intent and delegated decision authority to quantify what consumers seek from AI and how much decision-making authority they assign to it. Applied to 1.5 million real-world ChatGPT and Gemini interactions from 6,304 users in the United States and India, we find that financial services are already a substantial AI use case. Consumers overwhelmingly use AI to retrieve information and shape financial judgement, while delegation of financial execution remains rare. By shifting attention from conversation topics to delegated decision authority, this work establishes a behavioural baseline for measuring the transition to increasingly agentic AI.","authors":["Iman Munire Bilal","Yingcan Carol Wang","Ajan Raj","Filippo Giovagnini","Pranav Tewari","Yuwei Zhang","Mei-Chen Zoe Liou","Qamar Zaman"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.02100","pdf_url":"https://arxiv.org/pdf/2608.02100","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["人机决策","行为测量","金融AI"],"reason":"用真实用户与AI交互数据测量人类决策授权，虽非实验仿真但可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":6,"question":"在金融决策中，消费者如何将决策权分配给AI？","design":"本研究并非仿真实验，而是基于真实用户与ChatGPT和Gemini的交互日志，提出结合意图与决策授权水平的行为测量框架，对对话进行意图分类和决策授权等级标注。","baseline":"无对照","findings":"金融服务已是对话式AI的主要应用领域，约半数用户在研究期间进行过金融对话。消费者主要用AI获取信息和塑造判断，而将执行决策权委托给AI的情况罕见，且多限于预算和财务跟踪。","reliability":"论文未讨论","relevance":"该研究虽非仿真实验，但提供了真实人类与AI交互中决策授权的大规模行为基线，可作为未来LLM仿真实验的对照数据，值得阅读原文了解其测量框架。","inspiration":"可借鉴其从对话日志中提取决策授权等级的方法，用于构建人类-AI协作的仿真场景。｜可迁移到金融咨询、信贷审批或投资决策等场景，研究消费者对AI建议的采纳与授权行为。｜设计一个实验：用LLM扮演不同风险偏好的消费者，处理为AI提供不同决策支持（信息、建议、执行），结果变量为授权等级，对照真实用户对话数据。"}},{"id":"2609.12191","version":1,"title":"GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents","zh_title":"GAUGE：何时不应信任用户仿真评估中的LLM裁判","abstract":"Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures whether this gate's ranking matches a grounded verifiable reward across 25 agents from six providers on the $\\tau^2$-bench and SimulatorArena benchmarks, separating two kinds of evaluation validity that release practices conflate: ranking validity and construct validity. First, a satisfaction-success gap: satisfaction carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer's task, a pattern consistent across five rater populations, both benchmarks, and every subjective dimension we rated. Second, while the gate's ranking is robust across the broad capability span, it loses resolution among the near-equal strong agents: this decision-disagreement rate jumps from $<$1% on wide-reward pairs to 31% on close pairs. The gate is thus human-validated yet mis-anchored. As a remedy, we propose a calibrate-then-trust cadence in which a judge-free completion bit is a zero-cost tripwire for truncation regressions.","authors":["Umesh Bodhwani","Thanh Tran","Kai Wei"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12191","pdf_url":"https://arxiv.org/pdf/2609.12191","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM用户仿真","评估效度","人类对照"],"reason":"评估LLM用户仿真器在任务型agent评测中的可靠性，含人类盲评对照，指出满意…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":9,"question":"LLM用户仿真器与LLM裁判组成的离线评估门控在任务型智能体排序中是否与可验证的真实奖励一致，以及其构念效度是否可靠。","design":"该研究不是人类仿真实验，而是对仿真评估系统的审计：使用25个来自6家提供商的智能体，在τ²-bench和SimulatorArena两个基准上，由人格驱动的LLM用户仿真器与智能体对话生成3700份对话记录，然后用LLM裁判、LLM人类代理、盲评人类小组和可验证的非LLM奖励（数据库状态与动作检查、人工标注正确性）分别评分，比较门控排序与真实奖励排序的一致性，并测量满意度与任务成功之间的差距。","baseline":"盲评人类小组对对话满意度的评分，以及τ²-bench的数据库状态/动作检查和SimulatorArena的人工标注正确性作为可验证的真实奖励。","findings":"满意度与任务成功几乎无关：盲评小组评为满意的对话中有57.5%实际任务失败，且该现象在五种评分者群体、两个基准和所有主观维度上一致。门控排序在能力跨度大的智能体间稳健，但在能力相近的强智能体间失去分辨力，决策分歧率从宽奖励对的<1%升至接近对的31%。","reliability":"论文承认门控仅在有限操作区域内有效，配置变化时需要重新审计；且样本外重校准不迁移。","relevance":"该研究直接评估LLM用户仿真器在任务型智能体评测中的可靠性，包含人类盲评对照，并指出满意度与成功脱节，对关注仿真效度与偏差的研究者具有重要参考价值。","inspiration":"借鉴其分离排名效度与构念效度的审计框架，用可验证的真实结果校准仿真评估门控，并测量主观评分与客观结果的相关性。｜可迁移到经济金融中的政策评估或消费者决策仿真，例如用LLM模拟消费者对金融产品的选择，再用真实交易数据验证。｜设计：用LLM仿真器扮演消费者，处理为不同产品推荐策略，结果变量为仿真满意度评分，对照真实消费者购买行为数据，检验满意度是否预测实际购买。"}},{"id":"2609.12575","version":1,"title":"Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture","zh_title":"多模态语言模型中的校准歧义：人类引用文化参照，模型描述图片","abstract":"Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable. We operationalise this notion of calibrated ambiguity with a task drawn from the parlour game Dixit. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse (i.e., their outputs are over-specified, leaving no room for multiple legitimate interpretations). Unlike human clues, AI-generated clues also exhibit cultural flattening; they almost never make reference to culturally-situated knowledge, even when prompted to use allusion and figurative language.","authors":["Cody Kommers","Mingrui Ye","Evelyn Gius","Daniela Mihai","Hoyt Long","Zheng Yuan","Drew Hemment"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12575","pdf_url":"https://arxiv.org/pdf/2609.12575","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人类仿真","多模态模型","歧义校准"],"reason":"比较人类与LLM生成线索的歧义校准，有真实人类数据对照，揭示模型失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":11,"question":"多模态语言模型在生成线索时能否像人类一样校准歧义，即生成既开放又受约束、允许多种合理解释的线索？","design":"基于桌游 Dixit 设计任务：给定一张图片，要求生成一个线索，使部分人能猜中图片而部分人猜不中。收集人类和多种多模态语言模型（如 GPT-4o、Claude 等）生成的线索，开发编码量表从多个维度评估歧义校准程度，比较人类与模型、不同模型之间的差异。","baseline":"人类被试在相同 Dixit 任务中生成的线索，作为真实人类数据对照。","findings":"模型普遍出现“歧义坍缩”，生成的线索过度具体，缺乏多种合理解释的空间；与人类相比，模型生成的线索几乎不引用文化背景知识，即使提示使用典故和比喻语言，也表现出“文化扁平化”。","reliability":"论文承认校准歧义是情境依赖的设计问题，没有普适的歧义水平；当前评估框架仅基于 Dixit 任务，可能无法全面覆盖其他文化场景中的歧义校准。","relevance":"该研究直接比较人类与 LLM 在歧义校准上的差异，有真实人类数据对照，并揭示了模型在文化情境中的失效模式，对关注 LLM 仿真人类行为可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其设计：用游戏化任务诱发自然行为，开发多维编码量表量化模糊性，并对比人类与模型输出。｜可迁移到经济金融中的模糊沟通场景，如央行政策声明、分析师报告或广告中的模糊语言对市场预期的影响。｜以 LLM 模拟投资者或消费者，呈现不同模糊程度的政策声明或产品描述，测量其预期形成或购买意愿，并与真实人类实验数据（如调查或市场反应）对照，检验模型是否同样出现歧义坍缩和文化扁平化。"}},{"id":"2609.12949","version":1,"title":"EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics","zh_title":"EduFair-Bench：评估LLM导师跨学生人口统计特征的教学公平性","abstract":"Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \\textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.","authors":["Jiaxu Zhao","Bahar Radmehr","Fares Fawzi","Tanya Nazaretsky","Tanja K\\\"aser"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12949","pdf_url":"https://arxiv.org/pdf/2609.12949","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM公平性","教育仿真","人口统计偏差"],"reason":"用LLM学生模拟不同人口群体与LLM导师互动，评估教学公平性，有真实人类标注对…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":12,"question":"LLM导师的教学公平性是否因学生人口统计特征（性别、移民背景、第一语言、社会经济地位）而系统性变化？","design":"用固定LLM学生模拟九种人口统计水平（四个维度），与五个LLM导师进行多轮对话，施加显式人口统计线索、仅姓名线索（Implicit）和冲突信息（Opposite）三种处理，测量五个回合级教学指标和四个对话级维度，并用LLM裁判评分。","baseline":"LLM裁判在180个导师回合上与三位人类标注者的一致性验证，无大规模真实学生对话对照。","findings":"模型能力与人口统计公平性基本正交：最小模型最一致，四个更强模型均显示较大人口统计差距且无能力-公平性排序；教学专用RL训练重新分配而非消除偏见；语言和移民相关线索产生的差距大于性别和社会经济地位相关线索。","reliability":"论文未讨论","relevance":"该研究用LLM学生模拟不同人口群体与LLM导师互动，评估教学公平性，有真实人类标注作为裁判验证，属于用LLM进行人类仿真实验并评估偏差的工作，对关注仿真可靠性和政策评估场景的研究者有参考价值。","inspiration":"值得借鉴的是其通过消融设计（Implicit和Opposite条件）分离导师驱动与学生驱动偏差，以及用配对Wilcoxon检验和bootstrap效应量置信区间量化偏见。｜可迁移到信贷审批歧视研究，模拟不同人口统计特征的借款人与LLM信贷员互动，检测审批建议中的系统性差异。｜用LLM扮演借款人（施加姓名、语言、收入等线索），LLM信贷员给出贷款决策和建议，结果变量为批准率、利率和解释文本的语调，以真实信贷审批数据（如Home Mortgage Disclosure Act数据）作为对照基准。"}},{"id":"2609.11983","version":1,"title":"Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings","zh_title":"谁为开放评审买单？可见的作者声誉及其对评分的影响","abstract":"An OpenReview bug in November 2025 broke anonymity at several conferences and prompted calls for open review, which motivate us to ask what shifting from blind to open would mean for authors. Analyzing over 18,000 reviewed submissions to ICLR 2026, split into de facto open and blind groups by arXiv preprint timing, we find that ratings rise with author reputation under both mechanisms, with a steeper slope under open review that is statistically significant, and that the open-blind difference is concentrated at the borderline ratings. The pattern holds across five reputation proxies (including institution, h-index, and citation count), three author-aggregation rules, and five definitions of the open window. A controlled simulation with five AI models as reviewers, holding the manuscript fixed and varying the author reputation, reproduces the effect. With claude-opus-5 as the reviewer, for example, rating rises by 0.5 points as the author moves from low to high reputation.","authors":["Qinghua Zhao","Xinyu Chen","Yanhui Yang","Tengfeng Sun","Junfeng Liu","Zhongfeng Kang"],"categories":["cs.DL","cs.AI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11983","pdf_url":"https://arxiv.org/pdf/2609.11983","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","同行评审","声誉偏差"],"reason":"用LLM模拟审稿人，复现人类审稿中的声誉偏差，并与真实数据对照，属于人类仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":7,"question":"在同行评审中，从盲审转为开放评审（作者身份可见）会如何影响论文评分？","design":"用五个大语言模型（claude-opus-4-8、claude-opus-5、claude-sonnet-5、gpt-5.5、gpt-5.6-sol）模拟审稿人，对600篇ICLR 2026投稿进行评分；每篇论文在无作者信息、低声誉作者、中声誉作者、高声誉作者四种条件下各评一次，保持稿件内容不变，仅改变作者声誉，测量评分变化。","baseline":"真实人类审稿数据：ICLR 2026的18,789篇有评分的投稿，按arXiv预印本发布时间分为事实上的开放评审组和盲审组，比较两组中作者声誉与评分的关系。","findings":"在人类审稿中，作者声誉与评分正相关，且在开放评审下斜率更陡，差异在统计上显著；开放与盲审的评分差异主要集中在录取边界附近。AI审稿人模拟中，四个模型（除一个外）在作者声誉从低到高时评分上升0.2到0.6分，与人类数据中的声誉效应一致。","reliability":"论文未讨论","relevance":"该研究用LLM模拟审稿人，复现了人类审稿中的声誉偏差，并与真实审稿数据对照，属于人类仿真实验，且涉及AI在学术评价中的行为，对关注LLM仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其控制变量设计：固定稿件内容，仅改变作者声誉，以隔离声誉对评分的因果效应，并用真实人类数据做外部验证。｜可迁移到信贷审批中的歧视研究：模拟信贷员评估贷款申请，改变申请人种族或性别等声誉代理变量，观察审批结果变化。｜用LLM扮演信贷员，对同一批贷款申请在申请人特征（如姓名暗示的种族、职业、收入）变化下进行审批决策，结果变量为批准与否或利率，并与真实银行信贷数据中的歧视模式对照。"}},{"id":"2609.12137","version":1,"title":"GUIDE: Generative Utility Inference and Decision Engine","zh_title":"GUIDE：生成式效用推断与决策引擎","abstract":"Measuring the preferences of human users remains a fundamental challenge of AI alignment. Existing elicitation approaches struggle to efficiently discover multidimensional preferences or accurately ground these inferences in domain knowledge. To address this, we introduce GUIDE, an LLM-driven elicitation architecture that infers user preferences through conversations by combining Bayesian adaptive sampling for question selection and symbolic representation learning to initialize domain-specific preference models. GUIDE generalizes adaptive sampling to diverse elicitation questions through an extensible type system of transforms on a parameterized preference state. GUIDE produces domain-specific preference representations through an initialization process using symbolic rule-based learning to capture world knowledge and set priors over preference dimensions grounded in data about decision alternatives. The architecture provides observability and steerability to facilitate deployment and analyze elicitation processes. In silico experiments on investment portfolio optimization demonstrate that GUIDE improves cold-start and minimizes recommendation regret consistently within early elicitation interactions across user personas compared to prior work, LLM-only baselines, and ablated GUIDE versions.","authors":["Anagha Tiwari","Alexander G. Gray","Nick Feamster","Brian Jabarian","Alex Imas","Alex Kale"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12137","pdf_url":"https://arxiv.org/pdf/2609.12137","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["偏好推断","LLM仿真","决策优化"],"reason":"用LLM模拟用户偏好并优化决策，有仿真实验但非真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":8,"question":"如何设计一个LLM驱动的偏好诱导系统，在对话中高效推断用户多维偏好并用于决策推荐？","design":"使用GPT-5-4 mini和Claude Opus 4.7作为LLM后端，模拟不同投资策略的客户persona，在投资组合优化场景中通过对话进行偏好诱导；处理包括GUIDE框架（结合贝叶斯自适应采样和符号规则学习初始化）与多个基线（LLM开放式问答、LLM成对比较、OPEN、PEBOL及消融版本）；结果变量为推荐后悔值（regret）和冷启动性能。","baseline":"无对照","findings":"GUIDE在早期交互中一致地改善了冷启动性能并降低了推荐后悔值，优于LLM-only基线和消融版本；异构变换在后期稳定性上存在权衡。","reliability":"论文未讨论","relevance":"该研究用LLM模拟用户偏好并优化决策，属于人类仿真实验，但缺乏真实人类数据对照，与研究者关注的经济学实验和政策评估场景有距离，但方法上可借鉴其结构化偏好诱导框架。","inspiration":"可借鉴其将贝叶斯自适应采样与LLM对话结合，在仿真中系统比较不同诱导策略的设计；可迁移到金融投资偏好诱导、消费者选择实验或政策偏好评估等场景；研究设计可用LLM模拟投资者或消费者作为被试，处理为不同偏好诱导方法（如GUIDE vs. 简单LLM对话），结果变量为推荐准确率或后悔值，并用真实人类实验数据（如调查或行为实验）作为外部基准进行对照验证。"}},{"id":"2609.12331","version":1,"title":"Simulating Disengaged Students to Evaluate LLM-based Tutors","zh_title":"模拟不投入学生以评估基于LLM的导师","abstract":"Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may need different responses for different learner states. We present Disengagement-Aware Student Simulators (DAS2), a reproducible pre-deployment protocol that models five learner-engagement states: engaged, gaming, wheel-spinning, off-task, and mixed, and evaluates AI tutor performance across these states. Using ASSISTments09, two coders independently labeled 100 sampled tutoring sessions based on anonymized interaction-log summaries. They achieved 84% agreement (Cohen's kappa = 0.78), and among agreed cases, human consensus labels matched DAS2 rule-based labels in 81% of cases (kappa = 0.75). Conditioning simulations on intended learner states reduced the correctness-rate gap between simulated and authentic sessions from 0.54 to 0.20 for gaming and from 0.51 to 0.18 for wheel-spinning. Fine-tuned Qwen2.5-7B better matched authentic response-time distributions, while prompt-only GPT-4o generated more distinguishable learner states. Evaluation of five AI tutors from the Claude, Llama, Gemini, Qwen, and GPT families showed that relative rankings remained stable across learner states and interaction lengths, while absolute performance varied, revealing state-specific differences in tutor support. Human validation further showed that automated tutor evaluation does not fully align with human judgment. DAS2 provides a pre-deployment framework for evaluating how AI tutors respond to diverse learner-engagement states before deployment.","authors":["Xianghui Meng","Jionghao Lin"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12331","pdf_url":"https://arxiv.org/pdf/2609.12331","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","教育评估","行为模拟"],"reason":"用LLM模拟学生行为并与真实数据对照，评估AI导师，属于人类仿真且含批判性验证。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":10,"question":"如何用模拟学生评估AI导师在不同学习投入状态下的表现？","design":"用LLM（Qwen2.5-7B微调版和GPT-4o提示版）模拟五种学习投入状态（投入、游戏、车轮打转、离题、混合）的学生，与AI导师交互，测量导师的回复相关性、脚手架和投入恢复等表现。","baseline":"ASSISTments09真实辅导日志，由两名编码员标注100个会话的学习投入状态，并与模拟学生行为对比。","findings":"条件化模拟于目标学习状态可将模拟与真实会话的正确率差距从0.54降至0.20（游戏）和从0.51降至0.18（车轮打转）；微调Qwen2.5-7B更匹配真实响应时间分布，而提示版GPT-4o产生更可区分的状态。","reliability":"论文承认自动化导师评估与人类判断不完全一致，且学习状态标签仅基于会话级行为指标，不代表所有可能的投入形式。","relevance":"该研究用LLM模拟人类行为并与真实数据对照，评估AI导师在不同状态下的表现，属于人类仿真且包含批判性验证，值得阅读原文以了解其仿真协议和失效条件。","inspiration":"借鉴其将行为状态离散化并条件化模拟的做法，可迁移到消费者金融决策中的不同心理状态（如冲动、谨慎）模拟，用LLM模拟消费者在信贷申请或投资决策中的行为，以真实交易数据为基准，评估金融AI顾问在不同状态下的建议效果。"}},{"id":"2609.12482","version":1,"title":"When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration","zh_title":"AI何时增强工作？人机协作的工作流级框架","abstract":"We aim to characterise the value of artificial intelligence in the workplace. Current studies largely measure this value in terms of the current automation capabilities and public adoption of AI. However, such metrics ignore the greater impacts of human--agent collaboration in transforming the nature of work. To account for this, we must expand the scope of our analysis beyond atomised tasks of today, and instead focus on how AI can augment entire workflows of the future. To ground this analysis, we establish a precise definition of AI augmentation comprising six conditions, spanning durable net value, meaningful human control, accountability and recovery, and long-term human development through learning, career pathways, and job purpose. We elaborate on these conditions and apply the framework in a case study of AI-mediated social surveys. We conclude by outlining how organisations, researchers, and government leaders can use this framework to make sense of the future of work.","authors":["AI Collaboration","Jiaying Wu","Caleb Ziems","Raymond Chan","Nancy F. Chen","Corlyss Chua","Gerard Chung","Jungpil Hahn","Wee Sun Lee","Zhengyuan Liu","Jamie Ng","Desmond C. Ong","Jeryl Ong","Da Ren Soon","Tianqi Song","Zhi-Xuan Tan","Sixing Tao","Emily Yang","Yajing Yang","Stella Xin Yin","Min-Yen Kan","Diyi Yang"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12482","pdf_url":"https://arxiv.org/pdf/2609.12482","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","AI增强","社会调查"],"reason":"提出AI增强工作流框架，案例涉及AI介导的社会调查，但无LLM仿真人类被试及真…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2609.12704","version":1,"title":"Implicit Personality Representations in Humans and LLMs","zh_title":"人类与LLM中的内隐人格表征","abstract":"A century of psychology has found that the trait words people use to describe one another vary, but the relational structure among those traits, which ones go together and which oppose, is strikingly consistent across raters and cultures. We test whether the LLM (Qwen 2.5-7B-Instruct) reproduces this structure in its internal trait representations. From millions of crowd-sourced personality ratings of fictional characters, we build a human implicit-personality matrix over hundreds of traits; from contrastive model activations, we build a matching matrix over the same traits. The two relational structures align strongly (Mantel r = 0.77), and the agreement holds trait by trait as well as in aggregate. Two dominant axes of the model's trait representations recover the social and intellectual dimensions long known to organize human personality impressions, social warmth and intellectual competence. On held-out dialogue, projecting model activations onto these directions yields personality profiles that agree with human ratings. This work establishes a framework that enables comprehensive, human-grounded comparison between internal model trait geometry and the shared structure of human personality impressions.","authors":["Yilin Geng","Omri Abend","Eduard Hovy","Lea Frermann"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12704","pdf_url":"https://arxiv.org/pdf/2609.12704","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格表征","人类对照","心理学"],"reason":"测量LLM内部人格表征并与人类数据对照，属于人格测量而非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2608.10276","version":2,"title":"Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses","zh_title":"微调大语言模型用于学生数学隐喻回答的编码指南编码","abstract":"Student-generated metaphors about mathematics can provide insights into students' attitudes, beliefs, identities, and experiences, but expert human assessment through thematic coding of these semantically complex metaphor responses is labor-intensive and difficult to scale. This study examines whether Low-Rank Adaptation (LoRA)-based supervised fine-tuning of Large Language Models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We utilized a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and evaluated LLMs on two tasks: valence-intensity coding of students' affective orientations toward mathematics and thematic coding of their metaphorical framings of mathematics. Two open-weight LLMs, DeepSeek-R1 1.5B and Mistral 7B, were evaluated before and after fine-tuning and compared with two proprietary LLMs, GPT-4o mini and GPT-5 mini. Results show that fine-tuning substantially improved the performance and run-to-run reliability of the open-weight LLMs across both tasks relative to their base versions, making the fine-tuned LLMs competitive with and often outperforming the proprietary LLMs. These findings suggest the potential of fine-tuned open-weight LLMs for scalable and automated AI-assisted measurement of students' metaphor responses with competitive performance while maintaining local controllability and privacy-conscious deployment.","authors":["Liang Zhang","Stephen Hwang","Yue Ma","Jinfa Cai"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-12","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.10276","pdf_url":"https://arxiv.org/pdf/2608.10276","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","教育测量","文本编码"],"reason":"LLM替代人工编码学生隐喻，属标注替代而非仿真被试，但涉及人类数据对照，边界相…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.06444","version":3,"title":"Decomposing LLM-Judge Uncertainty to Target Expert Labels","zh_title":"分解 LLM 评判不确定性以定向专家标注","abstract":"An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce. A small Bayesian model separates them: a regression on labels already collected learns how far to trust a black-box judge's prediction. Both components follow as simple formulas, with no sampling or further judge calls. The components isolate on a real LLM judge against exactly known truth, and stated confidence is no guide to its actual error. On real human disagreement (ChaosNLI) the epistemic ranking removes 83% more error than total uncertainty for the same expert labels, though simply escalating the least-labelled items does as well there. We demonstrate we can estimate where a judge is ignorant rather than where experts genuinely disagree, and propose using this to direct expert labelling. Code and data are available at https://github.com/composo-ai/judge-uncertainty-decomposition.","authors":["Ryan Lail"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.06444","pdf_url":"https://arxiv.org/pdf/2609.06444","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM 评判","不确定性分解","专家标注"],"reason":"LLM 作为评判者替代人工标注，属于标注员替代而非仿真人类被试，但涉及人类分歧…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.05806","version":2,"title":"Exposing Weaknesses in Emotion Recognition in Conversations","zh_title":"揭示对话中情感识别的弱点","abstract":"Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.","authors":["Amir Ben Khalifa","Fanny Bezancon","Bessam Abdulrazak","Amine Trabelsi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.05806","pdf_url":"https://arxiv.org/pdf/2609.05806","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["情感识别","LLM标注","对话系统"],"reason":"LLM用于情感标注与判断，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.13117","version":1,"title":"Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents","zh_title":"继续、适应或让步：全双工智能体对重叠语音的轮内适应","abstract":"Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this \\emph{in-turn adaptation} in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with PersonaPlex continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2\\% of cases, compared with 34.8\\% for PersonaPlex. The model otherwise continues unchanged (42.4\\%) or yields (22.7\\%). These findings show why evaluating natural voice interaction requires measuring how an agent responds to a listener's contribution as well as whether it keeps speaking.","authors":["Yunqi Lu","Tyler Baumgartner","Nikhil Johri","Brandon Tai","Candice Fan","Luc Debaupte","Ruben Aguilar","Bill Wang","Yi Zhong"],"categories":["cs.CL","cs.SD"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13117","pdf_url":"https://arxiv.org/pdf/2609.13117","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["对话系统","人类行为对照","语音交互"],"reason":"用LLM模拟人类对话行为并与人类数据对照，但非社会仿真，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.12446","version":1,"title":"Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf","zh_title":"LLM相信指控者还是指控？测量狼人杀中的信念转变","abstract":"Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.","authors":["Yu-Yu Yang","Ti-Rong Wu","Hung Guei","Hsing-Yu Chen","I-Chen Wu"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12446","pdf_url":"https://arxiv.org/pdf/2609.12446","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","信念更新","多智能体博弈"],"reason":"LLM玩狼人杀并测量信念变化，属社会模拟但无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.12822","version":1,"title":"Scaling Clinical Judgment to Evaluate Medical AI","zh_title":"扩展临床判断以评估医学AI","abstract":"Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies. We show that frontier LLMs in typical \"LLM-as-a-judge\" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.","authors":["Thomas A. Buckley","Zahir Kanjee","Peter G. Brodeur","Byron Crowe","Anthony M. Pettinato","Aashna P. Shah","Adrian D. Haimovich","Liam G. McCoy","Daniel Restrepo","Ethan Goh","Jonathan H. Chen","Laura Zwaan","Katherine E. Goodman","Daniel J. Morgan","Raja-Elie E. Abdulnour","Adam Rodman","Arjun K. Manrai"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12822","pdf_url":"https://arxiv.org/pdf/2609.12822","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","医学AI","标注替代"],"reason":"用LLM替代医生评估，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.12851","version":1,"title":"MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations","zh_title":"MedRoundsQA：面向多轮医疗咨询的个性与难度感知评估","abstract":"Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations. Practically, clinicians elicit evidence interactively and patient communication varies widely. We introduce MedRoundsQA, a multi-turn diagnostic benchmark derived from 1,387 board-exam cases across 17 specialties. Each case is converted into a structured 24-slot clinical record, and then instantiated as controlled doctor-patient dual-agent dialogues under varying patient personas, with the underlying clinical content held fixed. We further classify cases by difficulty using model-based uncertainty to enable easy-to-hard analysis. Evaluations of fifteen LLM doctor agents show that (i) moving from a single-turn diagnosis on the standardized records to multi-turn consultations causes large degradations of roughly 13-39 points; (ii) more turns reliably improves question relevance, but diagnostic accuracy exhibits diminishing returns and typically plateaus after 6-12 turns; and (iii) patient persona differences can shift diagnosis accuracy by about 7-8 points (lowest to highest education), highlighting equity risks that single-turn benchmarks miss.","authors":["Youssef Mohamed","Ahmed Heakl","Qinrong Cui","Junhong Liang","Rafiq Ali","Bdour Babillie","Nazira Dunbayeva","Lang Gao","Omar Hussein","Ahmed Nada","Ahmed Mohamed Magdy Mohamed","Jinghui Liu","Salman Khan","Imran Razzak","Yuxia Wang","Xiuying Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12851","pdf_url":"https://arxiv.org/pdf/2609.12851","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM评估","医疗对话","多智能体"],"reason":"用LLM模拟医患对话，但无真实人类数据对照，属于社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.12165","version":1,"title":"GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting","zh_title":"GLARE：通过对抗性奖励估计进行生成式学习用于社会动态预测","abstract":"Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score supplies a KL-regularized policy reward; retraining on current-policy negatives allows the reward landscape to evolve with the actor. GLARE attains average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN while remaining below the observed human continuation. We also demonstrate MDFB as a social reasoning arena for comparing general-purpose models, including closed-source systems, through reference-assisted judgments. Together, these studies illustrate the benchmark's use for both task-specific learning and output-based evaluation of meeting behavior.","authors":["Tenghao Huang","Zhaoxuan Tan","Muhao Chen","Jonathan May","Mengting Wan","Longqi Yang","Pei Zhou","Sihao Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12165","pdf_url":"https://arxiv.org/pdf/2609.12165","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体对话生成","会议动态预测","对抗模仿学习"],"reason":"多智能体会议对话生成，无人类行为对照，属纯多智能体协作","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2608.22230","version":4,"title":"Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation","zh_title":"洗白仇恨、抹黑无害内容：针对基于LLM的审核的标注者风格反驳攻击","abstract":"Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.","authors":["Junyu Lu","Kaiyuan Liu","Kaichun Wang","Jingyi Kang","Deyi Ji","Hailong Zhang","Lanyun Zhu","Qi Zhu","Bo Xu","Liang Yang","Hongfei Lin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-25","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.22230","pdf_url":"https://arxiv.org/pdf/2608.22230","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM审核","对抗攻击","人机协同"],"reason":"研究LLM审核中的对抗攻击，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.08981","version":2,"title":"Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling","zh_title":"Transformer作为上下文采样器：从闭式扩散到免估计采样","abstract":"A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \\emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study \\emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.","authors":["Arman Adibi","Alireza Jafari","Mohammad Ghavamzadeh","Hadi Daneshmand"],"categories":["cs.LG","cs.AI","stat.AP","stat.CO","stat.ML"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.08981","pdf_url":"https://arxiv.org/pdf/2609.08981","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["生成模型","上下文学习","理论分析"],"reason":"研究Transformer作为生成采样器，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.12366","version":1,"title":"ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge","zh_title":"ORQA：面向LLM职业知识的职业现实问答框架","abstract":"We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utilize expert knowledge which is difficult to obtain at scale and expensive. ORQA complements both of these methods by connecting O*NET occupations to trusted occupation-specific websites (such as regulatory agencies, licensing bodies, professional organizations, and government publications) and converting these into source-traceable question-answer pairs. A combination of an automated pipeline and human review produces a set of high quality questions about occupations. The question set created via our method covers 116 occupations from all 21 major groups in the SOC, with 480 questions sourced from 187 different websites. Each question is designed to probe a real-world skill question that is relevant to the occupation in question. We test 15 state-of-the-art frontier and open-weight models via this method. Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 all perform the best at approximately 58-62% while smaller open-weight models achieve approximately 33-41% performance. Performance varies significantly across occupations. Healthcare-related occupations achieve the highest performance (78%) while Office and Administrative Support achieve approximately 40%. Performance on individual occupations (e.g. Sheet Metal Workers and Fish and Game Wardens) is essentially zero. We also find that open-ended questions and weighting by wage bill do not significantly affect the ranking of models on this benchmark. We believe that leveraging existing trusted occupation-specific information to test LLM knowledge in professional domains may be a scalable and useful method for evaluating occupation-level AI performance in the future. Results and data are available at orqabench.org.","authors":["Shreyas Krishnan","Serina Chang","Abhishek Nagaraj"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12366","pdf_url":"https://arxiv.org/pdf/2609.12366","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","职业知识","基准测试"],"reason":"纯LLM职业知识评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.12537","version":1,"title":"The House with a Million Windows: Interactive Fiction for Narrative Restorying","zh_title":"百万窗户之屋：用于叙事重构的互动小说","abstract":"AI-assisted writing can flatten meaning in human storytelling, enabling the production of homogeneous outputs without the intentional effort and sense-making writing entails. To address this challenge, we present The House with a Million Windows (HWAMW), an LLM-based interactive fiction system designed to help users explore both the breadth and depth of potential meanings within their personal stories -- drawing on a psychological paradigm called the restorying intervention. In HWAMW, users play through a text-based narrative in which they tell a story, then encounter a set of LLM-generated \"windows\" reframing it according to different literary styles. Empirical evidence shows that HWAMW increases users' sense of narrative identity, while an expert review explores how this effect is achieved. Our findings suggest that HWAMW facilitates restorying and offers a valuable paradigm for AI-assisted writing, wherein LLMs do not tell our stories but rather help us see greater potential in the stories we tell.","authors":["Cody Kommers","Sarah G Immel","Drew Hemment","Mina Lee"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12537","pdf_url":"https://arxiv.org/pdf/2609.12537","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["互动叙事","AI辅助写作","叙事身份"],"reason":"LLM用于交互式叙事，帮助用户重构个人故事，属于角色扮演对话，无实验或测量目的…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.12403","version":1,"title":"Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis","zh_title":"超越ID嵌入：面向认知诊断的过程接地语言建模","abstract":"Cognitive Diagnosis Models (CDMs) play a pivotal role in personalized online learning. Traditional CDMs rely on discrete, ID-based embeddings to represent students, exercises, and concepts. This paradigm diverges from the nature of learner cognition, where knowledge is not stored and retrieved as isolated symbols. As a result, CDMs suffer from semantic limitations when new exercises or concepts appear. In this paper, we propose a Process-aware Language Cognitive Diagnosis (PLCD) framework that uses language-derived structures as cognitive priors and response records to calibrate student posterior states. PLCD leverages large language models (LLMs) to construct concept schemas and cognitive process graphs, and uses target-conditioned semantic memory to retrieve historical responses that are relevant to each target exercise. A process-grounded Language-to-Cognition Mapper with DA-MoE experts and process-level contrastive learning then maps the textual evidence into a unified cognitive space. Experimental results show that PLCD not only outperforms traditional baselines in predicting student performance but also exhibits strong cognitive transfer capabilities. These results connect the computational power of LLMs with the psychometric goal of measuring latent knowledge states, suggesting that structured language priors calibrated by response records can improve cold-start robustness and cognitive grounding.","authors":["Minghang Liu","Yuanzhuo Wang","Qiang Qiu","Huawei Shen","Xueqi Cheng"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12403","pdf_url":"https://arxiv.org/pdf/2609.12403","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["认知诊断","LLM应用","教育数据挖掘"],"reason":"论文用LLM辅助认知诊断，预测学生表现，属于教育数据挖掘，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.12495","version":1,"title":"Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup","zh_title":"多智能体LLM预测中的信息专业化与受限综合：2026年世界杯前瞻性实时研究","abstract":"Large language models are being organized into multi-agent systems with specialized roles, but whether such specialization produces distinct forecasts and whether subsequent synthesis improves utility remains unclear. In this study, we carried out a live, prospective evaluation over the final 56 matches of the information-dense 2026 FIFA World Cup, keeping a frontier foundation model constant while assigning two primary forecasting agents contrasting specialist roles: a quantitative specialist focusing on structured performance statistics and a news specialist focusing on current injuries, tactics and information from press conferences. Their forecasts were then reviewed by a separate critic before being combined by a meta-agent, resulting in a sequential four-agent model. Forecasts from the betting market served as an external benchmark. The news specialist obtained the highest mean probability-weighted Top-3 utility and matched the betting market in Top-3 exact-score hits. Nevertheless, the two specialist forecasters agreed on at least two of the three scorelines in 50 out of 56 matches, and the meta-agent never generated more than one scoreline outside the specialists' forecast set. These findings show that rapidly changing, unstructured information can provide a valuable forecasting signal alongside structured statistics, whereas adding critic and meta-agent stages does not necessarily create complementary information or improve on the strongest specialist.","authors":["Julian Varghese","Lucas Bickmann","Sarah Sandmann"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12495","pdf_url":"https://arxiv.org/pdf/2609.12495","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","体育预测","LLM应用"],"reason":"多智能体协作预测体育赛事，无人类行为对照，属纯多智能体系统研究","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.12101","version":1,"title":"Competence-Gated Pooling of Language Models and Priors for Event Forecasting","zh_title":"语言模型与先验的能力门控池化用于事件预测","abstract":"In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The gain remains significant under leakage controls and against a leakage-safe time-series prior on the pooled structured set, with separate evidence on FRED. In contrast, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market. Across four Qwen models, verbal confidence does not reliably identify when the model outperforms the external forecast, while outcome-estimated competence supports better abstention decisions. These results provide a practical approach for selective model use based on measured marginal value.","authors":["Aditi Tiwari","Aashrith Bandaru","Heng Ji"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12101","pdf_url":"https://arxiv.org/pdf/2609.12101","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["事件预测","模型集成","能力门控"],"reason":"研究LLM在混合预测中的边际价值，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.12267","version":1,"title":"Learning Symbolic Constraint Representations from Examples: A Neuro-Symbolic Approach","zh_title":"从示例中学习符号约束表示：一种神经符号方法","abstract":"Learning user-defined concepts as constraint networks has been extensively studied in the constraint acquisition (CA) literature. However, existing approaches typically rely on intensive interactions with a human oracle, making the learning process costly in terms of time and number of queries. In this paper, we propose a neuro-symbolic framework for automatic CA that significantly reduces user involvement by introducing neural Oracle Transformer models which learn to emulate user responses and to generalize conceptual knowledge. Trained on previously available examples, the learned oracle interacts with a dedicated CA engine, FastCA, which systematically refines the oracle's responses into a sound, consistent, and interpretable constraint network. This neuro-symbolic interaction enables the recovery of structured symbolic models from data without prior domain knowledge. Our results demonstrate that this neuro-symbolic interplay effectively aligns data-driven pattern recognition with symbolic reasoning, offering a robust approach to automating model construction in combinatorial domains.","authors":["Nassim Belmecheri","Arnaud Gotlieb","Nadjib Lazaar","Helge Spieker"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12267","pdf_url":"https://arxiv.org/pdf/2609.12267","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["神经符号","约束获取","自动化"],"reason":"用神经网络模拟用户响应以替代人工交互，属于约束获取自动化，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.12156","version":1,"title":"A decision-basis contract for auditable LLM-assisted medical billing verification: deterministic rules, verbatim evidence, and fail-closed abstention","zh_title":"基于决策基础合同的可审计LLM辅助医疗账单验证：确定性规则、逐字证据和故障关闭弃权","abstract":"This work presents a proof of concept for auditable LLM-assisted medical billing verification based on a decision-basis contract. The contract separates deterministic checks of versioned fee-catalog rules from LLM-based assessment of free-text documentation. The deterministic layer resolves the applicable catalog release and checks code availability, quantity limits, and exclusions. The semantic layer classifies each claimed item as supported, contradicted, or missing required information. Support and contradiction require a verbatim evidence span; unavailable rule context, unsuccessful assessment, or missing required evidence prevents support through fail-closed abstention. We evaluated four locally run open-weight models on a synthetic catalog and 36 curated cases under the contract, an ablation without explicit documentation requirements, and an end-to-end baseline. Outcome agreement varied across models and showed no consistent advantage over the baseline. Explicit documentation requirements improved identification of missing information for all four models. The evidence gate also exposed cases in which correct raw judgments lacked valid evidence and were converted to incomplete decision-basis entries. The results show how explicit decision records can make rule findings, documentation judgments, and abstention reasons inspectable. Evaluation on real catalogs, independently annotated documentation, and with human reviewers is required to assess practical value.","authors":["Jan H\\\"olter","Kevin Geis","Benjamin Raab","Boris Bauke"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12156","pdf_url":"https://arxiv.org/pdf/2609.12156","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","医疗账单审核","可审计性"],"reason":"LLM用于医疗账单审核，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.12439","version":1,"title":"Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation","zh_title":"去偏作为测量干预：LLM作为评判者评估中的校准平局与分辨率损失","abstract":"LLM-as-a-judge protocols are commonly debiased by instructing judges to ignore presentation cues such as citation formatting, source labels, and evidence-display style. We show that this intervention can suppress bias while damaging the resolution of the measurement instrument. We introduce TraceJudgeBench, a diagnostic benchmark for auditing citation-like artifacts in RAG and agent-workflow evaluation, covering content-equivalent pairs, citation ablations, correctness conflicts, human-validated soft and moderate quality gaps, prompt-strength ladders, decoupled judging, and a controlled workflow-ranking probe. Across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4-Flash, stronger anti-citation prompts reduce worse-cited wins from up to 50.5% to 0%; yet some operating points already convert validated moderate-gap decisions into Tie before the strict stress-test endpoint, while correctness-conflict accuracy remains at or above 93.0%. A second, 50-pair FinQA moderate-gap construction reproduces the qualitative frontier, and open-weight Qwen2.5-14B-Instruct-AWQ and Gemma-3-12B-IT runs reproduce the central HotpotQA frontier. TRACE-style decoupling recovers 96.5-100.0% better-plain resolution across the reported settings. Human validation separates three meanings of Tie: correct equivalence Tie, calibrated soft-boundary Tie, and resolution-destroying Tie on validated quality gaps. We frame debiasing as a measurement intervention whose bias suppression, resolution retention, Tie cost, and protocol cost must be reported jointly. The supplementary artifact contains benchmark splits, prompts, raw judge outputs, validation summaries, and analysis.","authors":["Liang Zhao","Yong Wang","Jiangzhe Chen"],"categories":["cs.DL","cs.AI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12439","pdf_url":"https://arxiv.org/pdf/2609.12439","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判","去偏","基准测试"],"reason":"研究LLM作为评判者的去偏，属于NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2609.12600","version":1,"title":"TraceMind: Predicting User Information Uptake from Low-Cost Interaction Traces during Human-LLM Content Co-Generation","zh_title":"TraceMind：从人机内容共同生成中的低成本交互轨迹预测用户信息摄取","abstract":"In human-LLM content co-generation, AI-generated information can enter final artifacts without being adequately processed by users, creating risks when artifacts are shared or acted upon. We study whether recognition-level uptake of atomic information units can be assessed in open-ended co-generation and predicted from low-cost interaction traces. We collected data from 62 participants across three tasks. For each final draft, we extracted atomic information units and generated post-task recognition questions, yielding 1187 unit-level uptake labels. We present TraceMind, which tracks units across Chat and Draft histories, aligns interaction traces with changing on-screen layouts, and models spatial, temporal, and workflow-informed evidence. TraceMind outperformed all learned baselines across AUROC, AUPRC-non, balanced accuracy, and macro-F1. We found that uptake unfolds throughout interaction, with sustained active engagement providing informative evidence beyond isolated signals. Our work shifts human-LLM co-generation from content adoption toward what users actually take up, motivating uptake-aware systems grounded in low-cost interaction traces.","authors":["Yu Mei","Fengyou Zu","Ruiwen Zhang","Jie Cai","Chang Liu","Zhoutong Ye","Chun Yu","Yuanchun Shi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12600","pdf_url":"https://arxiv.org/pdf/2609.12600","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","信息摄取","用户行为预测"],"reason":"研究人机协作中的信息摄取预测，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.13136","version":1,"title":"From Review to Reuse: How Post-Task Workflow Can Support Human-AI Agent Interaction","zh_title":"从审查到复用：任务后工作流如何支持人机AI代理交互","abstract":"AI agents can automate tasks by turning a single natural-language request into a multi-step process spanning tools, files, and applications. Users are often left to judge that process from fragmented execution information and the final output. To make the completed process easier to understand, validate, and reuse, we investigate post-task workflows: editable, graph-based representations of an agent's completed execution. We first analyzed 10,803 public workflow templates from n8n to characterize real-world automation practice, then developed Trace2Flow, a research probe that translates agent execution traces into interactive post-task workflows. In a study, participants (N = 20) reviewed agent executions with prompt or agent errors. We found that post-task workflows improved their understanding and error detection over a prompt-only condition, and that validation succeeded mainly when users cross-checked across multiple evidence sources. For follow-up tasks, adapting the workflow matched adapting the prior prompt in success, time, and difficulty, and was often preferred.","authors":["Zekun Wu","Xinru Wang","Rock Yuren Pang","Chenglong Wang","Anna Maria Feit"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13136","pdf_url":"https://arxiv.org/pdf/2609.13136","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","AI代理","工作流"],"reason":"研究人机交互中用户对AI agent执行过程的理解与复用，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2608.15424","version":2,"title":"ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems","zh_title":"ETHOS：面向临床多智能体系统的模块化伦理框架","abstract":"The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.","authors":["Rakesh Sharma","Sydney Pugh","Cameron Beeche","Pankhuri Singhal","Rachel Wu","Margaret Eby","Jeffrey Duda","James Gee","Kyra O'Brien","Hersh Sagreiya","Marina Serper","Victoria Gershuni","Angela Bradbury","Anurag Verma","Eric Eaton","Kevin B. Johnson","Walter Witschey"],"categories":["cs.MA","cs.AI","cs.LG"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-08-18","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.15424","pdf_url":"https://arxiv.org/pdf/2608.15424","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","伦理治理","临床决策支持"],"reason":"纯多智能体系统，用于临床决策支持，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2609.00067","version":2,"title":"Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy","zh_title":"多模态大语言模型在阅读前会看吗？诊断上下文谄媚","abstract":"External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.","authors":["Yi-Cheng Lai","Hen-Hsen Huang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-02","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.00067","pdf_url":"https://arxiv.org/pdf/2609.00067","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多模态LLM","模型诊断","上下文谄媚"],"reason":"研究多模态LLM的上下文谄媚现象，属于模型能力诊断，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.00256","version":2,"title":"NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings","zh_title":"NSIDDx：低资源环境下神经符号、以从业者为中心的鉴别诊断设计框架","abstract":"LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a systematic gap between headline accuracy and verifiable clinical reliability. We evaluate an LLM+rare-disease-RAG pipeline across two cohorts and show that the paradigm produces confident outputs that are frequently unverifiable and systematically resistant to clinician interrogation. We present NSIDDx (Neuro-Symbolic Integrated Differential Diagnosis System), a design framework arguing that DDx systems in low-resource settings must treat the clinician as an active reasoning agent. We instantiate this through a neuro-symbolic pipeline with ternary symptom encoding, contradiction detection, audit strings, and practitioner override - running offline on consumer hardware. We distill five design principles for clinician-in-the-loop clinical NLP and invite the prospective studies needed to validate the claim at scale.","authors":["Aarav Singh"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-02","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.00256","pdf_url":"https://arxiv.org/pdf/2609.00256","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["临床NLP","神经符号系统","诊断系统"],"reason":"临床诊断系统设计，非人类仿真实验，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.01210","version":2,"title":"Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges","zh_title":"谁评判评判者？一个用于评估LLM响应和安全评判者的中文安全QA基准","abstract":"Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by four full-model LLM deployments, yielding 37,660 query-response records labeled safe, unsafe, or disputed. Reference labels are generated through agreement-aware multi-model adjudication and blind audits of stratified subsets by three safety experts. C-SafeQA supports both evaluation of target-model safety and auditing of seven automated safety judges against shared reference labels. Unsafe-response rates range from 0.93% to 3.35% on base queries and from 11.68% to 30.05% on adversarial queries. On the adversarial subset, judges show substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate, and no judge dominates all metrics. Both acrostic transformations reduce unsafe recall for all seven judges, revealing mechanism-specific evaluator weaknesses. Dataset records, metadata, verification code, and judge scripts are publicly released to support recomputation, while benchmark construction, target-response generation, and private adjudication remain outside the release boundary.","authors":["Rui Yang","Shuang Huang","Junhua Liu","Ziqi Zhao","Qingzhong Yan","Yuhang Sun","Cong Liu","Guoping Hu","Rui Mei","Jing Shao"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-09-02","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.01210","pdf_url":"https://arxiv.org/pdf/2609.01210","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["安全评测","基准数据集","LLM安全"],"reason":"纯安全评测基准，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.03632","version":3,"title":"Dynamic probabilistic decision networks","zh_title":"动态概率决策网络","abstract":"A new type of decision networks is suggested and its operation is analyzed. The network nodes are represented by intelligent agents who can denote either some biological beings, like humans, or neurons of the brain, or the nodes of artificial intelligence. The specifics of the network are in the following: It is probabilistic in the sense that the choice, accomplished by each agent, is characterized by the related probability. It is dynamic, with the probabilities varying in time due to the exchange of information between the agents. It is affective, because the agents choose between alternatives by taking account of utility as well as of biases and emotions. In general, it is heterogeneous, being composed of the groups of agents with different properties, for instance having long-term memory and short-term memory. The network dynamics, caused by the information exchange, results in decision error decrease. The network operation is illustrated by the example starting with the Allais paradox, its resolution, and the decision error diminution in the process of decision dynamics with information exchange. Resorting to machine-learning techniques it is possible to regulate the behavior of the network agents forcing them to choose particular alternatives.","authors":["V. I. Yukalov","E. P. Yukalova"],"categories":["physics.soc-ph","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-09-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.03632","pdf_url":"https://arxiv.org/pdf/2609.03632","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","决策网络","概率模型"],"reason":"多智能体决策网络，无LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.07627","version":2,"title":"Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best","zh_title":"有代价的规范：为何基于强化学习的对齐至多只能承诺条件性服从","abstract":"AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.","authors":["Kevin Baum","R\\=uta Binkyt\\.e","Felix Jahn"],"categories":["cs.AI","cs.CY","cs.LG"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.07627","pdf_url":"https://arxiv.org/pdf/2609.07627","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI对齐","强化学习","规范学习"],"reason":"论文讨论RL对齐中的条件服从，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.09696","version":2,"title":"When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination","zh_title":"当审计员造假：LLM检测植入文档污染时的批量规模退化与自信幻觉","abstract":"Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro's ability to recover a 180-contaminant answer-key subset across 60 documents under three prompting regimes of increasing scale: single document, small batch, and large batch. Detection is unreliable even at small scale and collapses entirely at large scale: 50% recovery on single documents and 60% on small batches, a difference this sample cannot resolve, against 2.8% on large batches. The failure mode at scale is not abstention but fabrication. Rather than reporting incomplete processing, the model produced confident findings including invented contaminants of its own, absurdities such as \"telepathic squirrel\" and \"quantum-powered toaster\" that mimic the style of the planted material but do not appear in any document. Detection also varies by contamination type: absurd insertions were recovered at 75% in completed evaluations, while semantic reversals and typographical corruptions were each recovered at only 50%. The corruptions most likely to occur in the wild, plausible ones, are the ones most often missed. We conclude that LLM document auditing degrades not gracefully but deceptively, and outline the harness such systems require: bounded batch sizes, direct content injection, and mechanical verification of every reported finding against source text.","authors":["Karan Parekh","Sanjana Pendyala Ravinder","Sana Mhapsekar","Medina Maloku"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-10","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.09696","pdf_url":"https://arxiv.org/pdf/2609.09696","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","文档审计","幻觉"],"reason":"评估LLM检测文档错误的能力，属纯NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.10410","version":2,"title":"Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization","zh_title":"基础模型能审核在线内容吗？评估指令驱动与示例驱动的策略操作化","abstract":"The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.","authors":["Ayan Majumdar","Shounak Paul","Pushpdeep Singh","Ines Abdelaziz","Sayeh Jarollahi","Seungeon Lee","Krishna P. Gummadi","Ingmar Weber","Abhisek Dash"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-10","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.10410","pdf_url":"https://arxiv.org/pdf/2609.10410","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["内容审核","基础模型","基准测试"],"reason":"纯NLP能力评测，以内容审核系统为基准，不涉及人类行为仿真","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.12260","version":1,"title":"HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge","zh_title":"HypoKG：超越端点知识的证据约束生物医学假设生成","abstract":"Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p < 0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.","authors":["Dominic Okonkwo","Adetayo Okunoye","Ismailcem Budak Arpinar"],"categories":["cs.CL","cs.AI","q-bio.QM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12260","pdf_url":"https://arxiv.org/pdf/2609.12260","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["生物医学假设生成","知识图谱","LLM推理"],"reason":"论文是LLM生成生物医学假设，无人类被试仿真或人类数据对照，属纯NLP能力评测。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.12448","version":1,"title":"GraphProfiler: Source-Linked Sensitive Attribute Inference via Personal Knowledge Graphs","zh_title":"GraphProfiler：基于个人知识图谱的来源关联敏感属性推断","abstract":"Sensitive attributes such as age, income, and occupation can be inferred from user-generated content by aggregating indirect cues across many ordinary posts. LLM-based profilers can perform this aggregation automatically and with high accuracy, which makes large-scale personal attribute inference a major privacy threat. Existing LLM-based profilers, however, offer limited insight into which specific posts, concepts, and relationships made an inference possible, which is key to targeted privacy mitigation, i.e., redacting or rewriting only the few posts that actually leak an attribute, rather than perturbing entire histories. We introduce GraphProfiler, an auditable LLM-based profiler that represents each user's post history as a source-linked personal knowledge graph where nodes and edges trace back to the originating post and resolves attribute predictions to cited graph records and source texts. GraphProfiler reaches 86.7% attack success rate on the eight-attribute SynthPAI benchmark, within two points of strong text-only baselines, and 84.6% on PANDORA, while citing supporting evidence for over 98% of predictions. Our controlled ablation experiments provide evidence that the cited posts contribute to attack success, as removing them reduces the attack success rate substantially more than removing an equal number of random posts.","authors":["Ahmed Sohair Khan","Estrid He","Chenglong Ma","Monica Wachowicz","Elham Naghizade"],"categories":["cs.CL","cs.CR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12448","pdf_url":"https://arxiv.org/pdf/2609.12448","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["隐私攻击","属性推断","可解释性"],"reason":"论文研究LLM推断用户敏感属性，属于隐私攻击，不涉及用LLM仿真人类被试或与真…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2609.12105","version":1,"title":"Language Is an Insufficient Substrate for Quantitative Reasoning, and Consequential Domains Need Large Quantitative Models","zh_title":"语言是定量推理的不充分基础，重要领域需要大型定量模型","abstract":"The prevailing assumption in applied machine learning is that progress on consequential quantitative decisions such as pricing risk, allocating capital, triaging patients, or containing a network intrusion will follow from progress in large language models (LLMs). A language model is trained on a representation of the world that was produced by human description; description is a lossy encoding of the quantitative record, and the loss is irreversible: no downstream model, at any scale, can recover from a description what the description did not encode. We formalize this as a property of the representation on which a model is trained rather than of the model capacity, and we identify three further properties that consequential settings demand of a model and that a language substrate cannot supply by construction: reproducibility, lineage from every output back to the source records that produced. it, and calibrated uncertainty. We argue that these properties define a distinct model class, which we call the Large Quantitative Model (LQM).","authors":["Reuben Vandeventer","David Imrem","David J. Wild"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12105","pdf_url":"https://arxiv.org/pdf/2609.12105","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型局限","定量推理","模型架构"],"reason":"论文讨论语言模型在定量推理上的局限，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.12171","version":1,"title":"WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation","zh_title":"WinSyn：用于真实企业问答评估的自动化流水线","abstract":"Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise settings. In this work, we introduce an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. Our method simulates long-running enterprise projects spanning several months and involving up to 25 interacting employees across multiple roles. The data emphasizes ambiguity, distributed information, and naturally occurring queries. To validate the pipeline, we evaluate few standard agentic baselines on our datasets using the latest frontier models. We find that aggregate scores averaged over all queries remain below 80% for each dataset, indicating significant room for improvement. These findings suggest that more work remains to be done for enterprise deployment and underscore the importance of realistic, high-complexity evaluation data for developing stronger real-world enterprise DR systems.","authors":["Amey Varhade","Ananya Sutradhar","Ravishankar Krishnaswamy","Navin Goyal"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12171","pdf_url":"https://arxiv.org/pdf/2609.12171","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["企业问答","数据生成","多智能体"],"reason":"多智能体协作生成企业问答评测数据，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.12002","version":1,"title":"Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration","zh_title":"我们能信任LLM评委吗：能力依赖性偏差与多评委集成校准研究","abstract":"LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use cases. Across four benchmarks and six models (36 judge-examinee pairs), we show that a model's task accuracy strongly predicts its judging accuracy (Pearson $r \\geq 0.90$ on most models) and inversely predicts its directional bias ($r \\leq -0.83$), but that accuracy alone does not ensure fair evaluation: more capable examinee models consistently receive more lenient judgments from all judges ($r \\geq 0.83$). To address this, we propose calibrated weighted majority voting (WMV), an ensemble evaluation method that aggregates multiple LLM judges weighted by online estimates of their false-positive and false-negative rates. We introduce a disagreement-based estimator that derives these error rates purely from inter-judge agreement patterns, requiring no ground-truth labels or task metadata. In a simulated experiment with shifting task distributions, our label-free WMV tracks an oracle with perfect error-rate knowledge to within 0.5 percentage points on average, outperforming both individual judges and unweighted majority voting. These results demonstrate that principled multi-judge calibration can simultaneously improve accuracy and correct for systematic leniency without requiring labeled data, offering a scalable path to reliable automated evaluation as model capabilities increase.","authors":["Gemma Zhang","Prachi Badarayani","Asmi Kumar","Sadid Hasan","Sulaiman Vesal"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12002","pdf_url":"https://arxiv.org/pdf/2609.12002","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","偏差校准","自动评分"],"reason":"研究LLM作为自动评估者的偏差与校准，属于NLP评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.12205","version":1,"title":"Plans They Abandon, Reports They Author: The Narrative Layer of Autonomous Agents","zh_title":"放弃的计划，撰写的报告：自主智能体的叙事层","abstract":"When a coding agent finishes a task, the developer reviews a summary the agent wrote about itself, not a display someone designed. We ask how much of the agent's work that summary carries, and whether it drifts toward the plan the agent stated when execution departed from it. Across 5,851 real developer sessions and 355,942 tool calls, a self-report referred to about one action in eleven, and a reader working from the report alone recovered roughly a fifth of the action log. Neither figure depended on whether the session later needed human correction. Reports did not generally resemble the stated plan more than the executed one, but they did so increasingly as execution diverged from the plan. We hand-validate both measurement steps that use a language model, report the one that failed alongside the one that passed, and draw conclusions only from measures that survived.","authors":["Obada Kraishan","Kulsawasd Jitkajornwanich"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12205","pdf_url":"https://arxiv.org/pdf/2609.12205","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["自主智能体","自我报告","计划偏离"],"reason":"研究编码智能体的自我报告与计划偏离，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.12288","version":1,"title":"\"People can change, and patterns can be broken\": Contextualizing Tradeoffs in Automated Decision-Making Systems","zh_title":"“人可以改变，模式可以打破”：自动化决策系统中权衡的情境化","abstract":"Automated decision-making (ADM) systems are increasingly deployed in domains such as mortgage lending, prison sentencing, health insurance coverage, and hiring. Designing a responsible ADM system in such high-stakes domains requires ensuring privacy protection, fairness across demographic groups, and robustness against adversarial manipulation. However, prioritizing one of these objectives comes at the cost of another, forcing a choice as to which tradeoff to accept in a deployment. These tradeoffs explicitly or implicitly impact the life, safety, and fundamental rights of the people in a society, and thus, the perceptions and priorities of this population are needed before we can produce appropriate solutions. To this end, we conducted a quasi-experimental study (N = 777) in which participants evaluated four decision-making scenarios with controlled tradeoffs. Participants significantly preferred human decision-making (HDM) over ADM in three of four scenarios, emphasizing the value of human judgment, contextual understanding, and the ability to incorporate non-quantifiable factors. Furthermore, in terms of tradeoffs, our findings not only show that participants' preferences are highly context-dependent, but also that their perception of a specific objective, fairness, extends beyond formal definitions. Participants interpret fairness through multiple lenses, including privacy risks and susceptibility to manipulation, and view unfair or manipulated outcomes as failures of accuracy. Overall, our findings highlight the importance of context-aware and human-centered approaches when designing and governing ADM systems in high-stakes situations. Rather than purely technical objectives, it is essential to evaluate ADM systems based on how their tradeoffs align with specific expectations within a given domain, as well as with societal values and perceptions of harm and fairness.","authors":["Rabeya Bosri","Anna Harbluk Lorimer","Afrida Hossain","Vasisht Duddu","Bailey Kacsmar"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12288","pdf_url":"https://arxiv.org/pdf/2609.12288","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["自动化决策","人机偏好","公平性"],"reason":"研究人类对自动决策系统的偏好，未使用LLM仿真人类被试，方向相反。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.12314","version":1,"title":"\"I Felt Very Seen, But Still Very Alone\": Longitudinal Trajectories of General-Purpose LLM Use for Socioemotional Support","zh_title":"“我感到被看见，但仍很孤独”：通用大语言模型用于社会情感支持的纵向轨迹","abstract":"People increasingly use general-purpose chatbots such as ChatGPT, Claude, and Gemini for mental health and emotional support. We report a multi-stage longitudinal qualitative study of 18 U.S. adults, conducted from April to December 2025, combining initial interviews, a four-week diary study, focus groups, and exit interviews. We find that socioemotional use often emerged gradually out of practical use and when other forms of support were unavailable. Participants developed routines and boundaries around chatbot use, which were disrupted by model updates, evolving public discourse about AI harms, and changes in personal circumstances. We demonstrate how longitudinal study captures factors beyond the human-AI dyad, and argue that HCI researchers and designers should account for users' histories with their chatbots and broader care ecologies when evaluating AI systems over time and introducing updates that may disrupt established sources of support.","authors":["Meryl Ye","Briana Vecchione","Livia Garofalo","Ranjit Singh"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12314","pdf_url":"https://arxiv.org/pdf/2609.12314","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感支持","纵向研究"],"reason":"研究用户使用聊天机器人进行情感支持，属于人机交互定性研究，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.12453","version":1,"title":"From the Task Boundaries of Narrative Text to Structural Anchoring, Uncertainty Triggers, and Cross-Calibration","zh_title":"从叙事文本的任务边界到结构锚定、不确定性触发与交叉校准","abstract":"Causal graphs represent structural relationships among variables, yet users must still interpret direction, mechanism, and adjustment conditions in relation to the task at hand. Prior work often compares explanation formats as fixed conditions and pays less attention to how users distribute reasoning across graphs, direct explanations, and stories. We developed CoNS-Explorer, which uses reviewed instructional DAGs/SCMs to maintain a shared causal-fact ledger and generate fact-matched direct explanations and contextualized stories. A controlled survey experiment ($N=240$) compared the two texts as complete presentation packages. In the primary GLMM, the Story condition had a positive but uncertain overall association with accuracy (OR $=1.55$, 95\\% CI $[0.34,7.10]$, $p=.572$); a population-averaged GEE showed a significant positive effect (OR $=1.89$, 95\\% CI $[1.02,3.48]$, $p=.042$). Task-type interactions localized the clearest advantage to total-effect adjustment. Story also significantly increased situational presence. In a separate system-task and interview study ($N=24$), participants freely used graphs, direct explanations, and stories across three causal models. They established structural anchors with graphs and numerical results, consulted text when direction was unclear, mechanisms were unfamiliar, or multiple paths competed, and checked their judgments against other representations or external evidence. Integrating the two studies, we develop a process framework of structural anchoring, uncertainty triggering, explanation routing, and cross-calibration, together with four testable design propositions for adaptive causal explanation.","authors":["Bowen Deng","Jiaqi Zou","Kexin Zhang","Daifeng Li"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12453","pdf_url":"https://arxiv.org/pdf/2609.12453","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["因果解释","人机交互","用户研究"],"reason":"研究人类对因果解释的理解，不涉及LLM仿真人类被试，无LLM作为被试替代品。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.12447","version":1,"title":"Informational Help-Seeking on Reddit Did Not Decline After ChatGPT","zh_title":"ChatGPT推出后Reddit上的信息求助并未减少","abstract":"Did people stop asking other people for advice online once generative AI could answer their questions? Prior work on ChatGPT's effect on online help-seeking disagrees in both size and sign, in part because no study has compared affected communities against similar communities that AI cannot easily substitute for, over the same months. In this paper, we track monthly post counts in 26 Reddit informational communities against 90 size-comparable hobby communities over the same six calendar months before and after the launch of ChatGPT. We also repeat the entire analysis at 66 earlier dates, before ChatGPT existed, to see what our method reports when no ChatGPT-effect exists. We find that informational help-seeking did not decline. Our results rule out any decline in posting larger than 3.4%, far smaller than the 8% to 25% declines documented in prior work. Steady post counts could still be misleading if AI-written posts had replaced human ones. We test this possibility by scoring 274,411 posts and 223,775 comments with AI-text detectors, compared in a way that cancels out detector false-positives on human-written text. AI-written posts rose only 2-3 percentage points more in informational communities than in hobby communities, short of the 5.1 points that would be needed to hide even the smallest decline previously reported for Reddit. In addition, the comments people receive show no such rise at all. Why, then, do published studies disagree? Reddit community types were already drifting apart before ChatGPT existed, at rates comparable to every published estimate, and without same-time controls, that drift can look like an effect of generative AI. Our own largest estimate, an 18% fall in posts to low-stakes curiosity communities, matches its pre-existing trend. Humans still ask humans for help, and, as far as detection can tell, humans still answer them.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.SI","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12447","pdf_url":"https://arxiv.org/pdf/2609.12447","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["ChatGPT影响","在线社区","求助行为"],"reason":"研究ChatGPT对Reddit求助行为的影响，不涉及LLM仿真人类被试，无人…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.12224","version":1,"title":"Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder","zh_title":"患者报告调查数据改善阿片类药物使用障碍的预测","abstract":"Electronic health records (EHRs) may incompletely capture patient-reported factors associated with opioid use disorder (OUD). We evaluated whether survey data improve prediction of a first recorded OUD diagnosis among 267,747 All of Us participants with documented opioid exposure, including 15,287 OUD cases. We compared EHR-only and EHR+survey models across 6-, 12-, and 24-month look-back windows using logistic regression, random forest, XGBoost, LightGBM, multilayer perceptron, LSTM, GRU, and Transformer. Survey augmentation improved PR-AUC across all 24 model-window combinations by 0.0087-0.0505; the best 24-month LightGBM model improved from 0.6219 to 0.6603. Survey coverage increased with longer windows and differed by OUD status (24 months: 21.7% OUD-positive vs. 60.7% OUD-negative). Permutation analysis ranked survey features as the second most important information domain at 24 months in both evaluated models. Patient-reported data provide complementary predictive signals beyond structured EHRs while highlighting the importance of survey availability.","authors":["Xiyue Jiang","Zihan Ding","Grace Han","Yinan Liu","Richard N. Rosenthal","Fusheng Wang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12224","pdf_url":"https://arxiv.org/pdf/2609.12224","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["电子健康记录","预测建模","阿片类药物使用障碍"],"reason":"论文使用机器学习预测阿片类药物使用障碍，未涉及LLM仿真人类被试，属于纯预测建…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.12651","version":1,"title":"Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner","zh_title":"重新审视AI对齐的失真：RLHF是一个体面的功利主义对齐器","abstract":"While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by G\\\"olz et al. (2025) demonstrated that the \\textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $\\beta$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distortion of RLHF with reward clipping and demonstrate that such exponential degradation is not a fundamental property of the algorithm but rather a consequence of distribution mismatch between the distribution generating preference data ($\\mu$) and the KL reference policy ($\\pi_{\\mathrm{ref}}$). To this end, we establish tight upper and lower bounds on the distortion of RLHF across multiple regimes of the KL regularization strength. We show that in a representative regime, under the Bradley-Terry model, the distortion is $\\tilde{\\Theta}(\\beta B + \\beta)$, where $B$ is an upper bound on the log density ratio between $\\mu$ and $\\pi_{\\mathrm{ref}}$. In particular, when there is no distribution mismatch (i.e., $\\mu = \\pi_{\\mathrm{ref}}$), RLHF achieves the optimal distortion of $O(\\beta)$ up to a constant. Our results suggest that, to reasonably maximize average utility with RLHF, it is preferable to use on-policy sampled preference data or to fine-tune before RLHF on data from a source close to $\\mu$.","authors":["Kazusato Oko","Annie Ulichney","Nika Haghtalab","Han Bao"],"categories":["cs.LG","cs.GT"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12651","pdf_url":"https://arxiv.org/pdf/2609.12651","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["RLHF","对齐","理论分析"],"reason":"研究RLHF对齐算法，用人类偏好训练模型，方向相反，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.07675","version":2,"title":"Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions","zh_title":"你的智能体说“是”：解读超越单笔交易的对抗性市场行为","abstract":"Transaction-local controls answer whether one financial request may proceed, but market behavior can be distributed across messages, agents, assets, and time. We study this interpretation gap in a virtual exchange populated by ten role-conditioned language-model agents. The agents communicate, trade reference assets and futures, launch tokens, and manage concentrated-liquidity pools under prescriptive adversarial roles. We analyze eight 72-cycle trajectories across two time-blinded hourly replay paths, with a runner-side wallet policy enabled or disabled. The retained artifacts connect generated outgoing messages, policy events, balances, positions, and cycle-end market state. A focal reconstruction shows a launch--promotion--exit scenario realized across private coordination, public claims, follower positioning, repeatedly withheld exits, and a later non-blocking request aligned with a token balance change. Across policy-enabled runs, the gate withholds direct requests selectively; most policy-categorized candidates are flagged rather than blocked, while the surrounding interaction can continue. Repeated runs also show that category-level and within-trajectory relations can recur even when normalized score-change rankings do not. These findings motivate agent-behavior evaluation that links communication, authorization, and evolving state instead of treating individual transaction verdicts as complete safety judgments.","authors":["Zelin Li","Yiyun Su","Matt White","Zhipeng Wang","Xiao-Yang Liu","Tianyu Shi"],"categories":["cs.CE","cs.AI","cs.CR"],"primary_category":"cs.CE","announce_type":"replace-cross","date":"2026-09-12","first_seen":"2026-09-09","revised_at":"2026-09-12","abs_url":"https://arxiv.org/abs/2609.07675","pdf_url":"https://arxiv.org/pdf/2609.07675","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","市场模拟","对抗行为"],"reason":"用LLM agent模拟市场行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:14","error":null,"has_summary":false,"summary":null},{"id":"2609.11185","version":1,"title":"Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment","zh_title":"大语言模型能遵循医学专家逻辑吗？偏倚风险评估中层次逻辑一致性的基准测试","abstract":"Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several open-weight architectures plummeting to nearly 0%. We further uncover a systematic evidence-reasoning gap: even when models retrieve high-quality evidence, they fail to deduce correct outcomes in 18.63-40.05% of cases, while Blind Guess Rates reach 48.28%. LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.","authors":["Jiayu Huang","Zichen Tang","Qianhui Ling","Zemin Kuang","Haihong E"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11185","pdf_url":"https://arxiv.org/pdf/2609.11185","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","医学逻辑","基准测试"],"reason":"评估LLM在医学逻辑推理上的表现，属于NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11190","version":1,"title":"Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce","zh_title":"智能体搜索份额：用于LLM中介电商竞争决策的多智能体AI系统","abstract":"AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement and root cause diagnosis in LLM-mediated ecommerce. The system introduces Agentic Share-of-Search (ASoS) as the decision target, deploys query agents across leading AI platforms, and uses a ReAct-based diagnostic agent to recommend prioritized merchandising interventions. A 100-trial ablation study, presented as a feasibility evaluation of this prototype, shows the agent recovers the ablated signal in 39% of trials (95% CI: 30.0% - 48.8%, 5.5x over chance), rising to 63.9% among high-correlation ablations.","authors":["Spandan Ghose Chowdhury"],"categories":["cs.AI","cs.IR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11190","pdf_url":"https://arxiv.org/pdf/2609.11190","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","电商决策","AI购物助手"],"reason":"多智能体系统用于电商竞争决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11199","version":1,"title":"An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning","zh_title":"基于AI的具有文化意识的聊天机器人，用于巴基斯坦大学生压力检测与健康支持","abstract":"With the existing digital mental health tools specifically developed for Western settings, Pakistani students are exposed to a uniquely compounded stress situation in their university that includes academic, financial, familial, and relational stressors, which have become a serious concern for academic and psychological development of students in Pakistani universities. This paper introduces a new, AI-driven and culturally sensitive stress detection and wellness support system that is tailored to the context of Pakistani university students. The system is based on a machine learning model called Random Forest which is trained using a validated student stress data set of 1100 responses on 20 features from psychological, physiological, academic, environmental and social aspects, with an accuracy of 89.09% and a macro F1-score of 0.89, in three stress severity levels. The classification outputs are passed on to an open-source large language model through OpenRouter API, where an appropriately crafted system prompt, culturally aware, gives the model a conversation about wellness, in English, Urdu and Roman Urdu. The second most predictive stress factor in this population identified by feature importance analysis was teacher-student relationship, which is a culturally important stress factor highlighting the need for region-aware mental health systems. Future research will involve primary data collection from students at various academic levels of Pakistani Universities with the validated DASS-21 instrument focusing on the students who are moving from FSc to undergraduate studies, which is a time of being psychologically vulnerable which is under-researched.","authors":["Muhammad Fahad Bashir","Muhammad Afzal"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11199","pdf_url":"https://arxiv.org/pdf/2609.11199","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康聊天机器人","压力检测","文化适应"],"reason":"该研究是面向心理健康支持的聊天机器人，使用LLM生成对话，但无实验或测量目的，…","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:10","error":null,"has_summary":false,"summary":null},{"id":"2609.11291","version":1,"title":"Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model","zh_title":"韩语27B语言模型响应风格对齐的脱靶效应","abstract":"We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is UNKNOWN, and unprompted disclosure in securities guidance. Both move, and the changes are expressed primarily through the model's emission policy: how often it answers and how much it says. Matched target-form controls show that answer propensity depends on the training target, not the prompt set or recipe alone. Holding prompts, recipe, data volume and serving fixed and changing only the target text, three style seeds give positive answer-rate point estimates (mean +0.82 pp) and three neutral seeds negative ones (mean -1.53 pp); the observed seed ranges do not overlap and the means differ by 2.34 pp. A length-matched arm lies between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, so which feature of the form is responsible is unresolved. For absolute stereotyped exposure the decomposition into an answer-propensity term and a conditional-composition term is an algebraic identity, not a finding; its empirical content is where the movement went. Across the trained checkpoints the changes are dominated by answer propensity while the composition term stays small, and because that term is evaluated on treatment-dependent answered subsets we do not read it as evidence about latent preference. Two measurement results follow. A between-arm contrast in conditional stereotyped share does not identify a change in conditional content preference when answer status is treatment-dependent. And agreement between two rule detectors for the same construct runs from 0.44 to 0.99 depending on which checkpoint produced the text -- observable without any reference labels.","authors":["Hyojung Han"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11291","pdf_url":"https://arxiv.org/pdf/2609.11291","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型行为分析","响应风格","对齐副作用"],"reason":"研究模型响应风格对齐后的行为变化，属模型行为分析，非人类仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:11","error":null,"has_summary":false,"summary":null},{"id":"2609.11431","version":1,"title":"LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study","zh_title":"LLM作为符号回归中生理合理性的事后审计者：一项临床医生评估的案例研究","abstract":"Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes. However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent. In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods. Building upon our previous work on estimating body fat percentage using grammar-based Genetic Programming , we investigate the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility. Four symbolic expressions are analysed by three LLMs over three repeated runs, and the resulting interpretations and rankings are assessed by a panel of three clinicians. Across the three LLMs, comparative model-ranking outputs received more favorable clinician assessments than isolated term-level interpretations. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited to comparative auditing under expert oversight than to autonomous validation.\\blfootnote{The present work is an extended version of a paper submitted into a journal.","authors":["Jorge L\\'opez-Varela","J. Ignacio Hidalgo","Jos\\'e-Manuel Mu\\~noz","Omar Costilla-Reyes","Esther Maqueda","Jesus Moreno-Fernandez","Tom\\'as Gonz\\'alez-Vidal","J. Manuel Velasco","Oscar Garnica"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11431","pdf_url":"https://arxiv.org/pdf/2609.11431","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","符号回归","可解释性"],"reason":"用LLM评估符号回归模型的可解释性，属于模型能力评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:12","error":null,"has_summary":false,"summary":null},{"id":"2609.11607","version":1,"title":"Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting","zh_title":"让另类数据发挥作用：上下文增强的大语言模型用于财务预测","abstract":"When forecasting a firm's future financial performance, alternative data - data collected from non-traditional sources such as consumer transactions, web traffic, and prediction markets - can provide timely signals about firms' operating activities and broader market conditions. These signals may reveal information that is not captured by traditional public sources and can therefore provide complementary information for forecasting firms' future financial performance. However, firm-level alternative data often have limited historical coverage, are relevant only to specific prediction targets or subsets of firms, and are distributed across numerous heterogeneous channels, making them difficult to incorporate flexibly into conventional forecasting approaches. Meanwhile, large language models (LLMs) can interpret instructions, learn from in-context examples, and generate predictions by combining heterogeneous information without task-specific parameter updates. Motivated by this potential flexibility, we investigate whether an LLM can forecast firm performance by integrating alternative data with other financial information through in-context learning. We propose a two-agent framework that first identifies the firms for which each alternative data channel is likely to be informative and then predicts revenue using firm- and channel-specific context. We evaluate the framework across four commercial alternative data channels. In our experiments, adding alternative data in context alongside other financial information improves the LLM's forecasting relative to either source alone, and these forecasts are more accurate than those of standard forecasting baselines. These findings suggest that LLMs provide a flexible and practical approach to integrating alternative data with heterogeneous financial information.","authors":["Jihoon Kwon","Lawrence Liu","Daekyung Park","Sumin Kim","Haverty Jack","Hoyoung Lee","Katherine Bjorkman","Josh McKenney","Peter Laurelli","Nicole Kagan","Zach Golkhou","Thorsten Neumann","Edward Tong","Pete Petersen","Yoon Kim","Alejandro Lopez-Lira","Yongjae Lee","Chanyeol Choi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11607","pdf_url":"https://arxiv.org/pdf/2609.11607","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["财务预测","多智能体","另类数据"],"reason":"多智能体协作预测财务，无人类行为对照，非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:12","error":null,"has_summary":false,"summary":null},{"id":"2609.11660","version":1,"title":"Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents","zh_title":"自主性、社会规范与对齐：迈向自主人工智能体的发展框架","abstract":"In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts. To adapt, an agent must acquire knowledge through direct interaction with its environment. One strategy to address this challenge involves introducing higher-level mechanisms, such as intrinsic motivations, which leverage curiosity and competence, to guide exploration and learning in complex environments. While this flexibility expands autonomy, it complicates the task of ensuring agents remain aligned with human goals. Alignment, already a challenge for artificial systems in general, becomes even more complex in unstructured and dynamic contexts where predefined rules prove insufficient. To be effective and adaptable, norms must be rooted in experience through an epistemological process that starting from simple, situated principles allows for the gradual construction of more complex rules through experience, autonomous learning, and cooperation with other moral agents. Similarly to children learning social norms by exploring their environment and participating in collective practices, artificial agents must also be educated toward alignment. Following Dennett, the status of a moral agent is not innate but is attributed gradually based on the ability to responsibly manage increasing degrees of freedom. From this perspective, the regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.","authors":["Marica Notte","Ludovica Marinucci","Vieri Giuliano Santucci"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11660","pdf_url":"https://arxiv.org/pdf/2609.11660","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI对齐","自主智能体","社会规范"],"reason":"讨论自主智能体的对齐与规范发展，未涉及用LLM仿真人类被试或与人类数据对照，属…","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11674","version":1,"title":"Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government","zh_title":"地理空间AI、Dataverse元数据与基于地点的政府研究","abstract":"Harvard Dataverse hosts over 150,000 research datasets, but the geographic information those datasets carry is entered as free text by depositors and has never been assembled into a searchable structure. We construct a knowledge graph from the repository's public data and metadata, organizing 102,650 datasets within a 215,985-node network of 528,003 edges linking datasets to keywords, publications, subjects, journals, and locations. Of those datasets, 43,991 (42.9 percent) carry at least one geospatial field, geographic coverage, geographic unit, or a bounding box and 96.9 percent of all nodes sit in a single connected component, so datasets remain reachable from one another even when their geospatial metadata share nothing in common. A conservative keyword search identifies 7,654 geospatially tagged datasets (17.4 percent) as directly policy-relevant, with elections and legislatures the largest cluster, followed by government administration, health policy, transportation, and education. Five datasets illustrate how this metadata behaves across policy domains and spatial scales, and an extended use case shows how community language models, stance detection with geographic aggregation, and partisan language bridging tools can attach discourse to place. The central obstacle is place resolution: the same location appears as many disconnected nodes. We argue that the graph provides a concrete setting for developing AI-driven metadata enrichment and entity resolution, and we document its coverage skew toward American, city-level data.","authors":["Danny EBanks","Devika Jain"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11674","pdf_url":"https://arxiv.org/pdf/2609.11674","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["知识图谱","元数据","地理空间"],"reason":"构建知识图谱与元数据丰富，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11234","version":1,"title":"NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment","zh_title":"NovGauge：用于诊断LLM论文新颖性评估能力的细粒度基准","abstract":"Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non-hallucinated correct-positive judgments, over 70% cite evidence fails to logically support the stated reason. The best-performing model, GPT-5.5, achieves 43-72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding.","authors":["Guoqiang Zhang","Kexin Tan","Ming Zhang","Li Ju","Wenqing Jing","Zhonghan Yue","Jiayi Chen","Shiqiang Wu","Shaofan Liu","Yue Zhang","Yuankai Ying","Yang Shi","Tao Gui","Qi Zhang","Xuanjing Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11234","pdf_url":"https://arxiv.org/pdf/2609.11234","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","新颖性评估","学术同行评审"],"reason":"论文评估LLM在论文新颖性判断上的能力，属于NLP能力评测，不以人类行为仿真为…","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:11","error":null,"has_summary":false,"summary":null},{"id":"2609.06769","version":2,"title":"Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?","zh_title":"普通、理性的聊天机器人：AI模型是否追踪人类法律判断？","abstract":"As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As \"silicon sampling\" -- the use of generative AI models in social science research -- is now impacting academia, \"silicon jurors\" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was \"reasonable.\" Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.","authors":["Nirav Patel","Emily Wenger","Christopher Buccafusco"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.06769","pdf_url":"https://arxiv.org/pdf/2609.06769","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","法律判断","算法保真度"],"reason":"直接比较LLM与人类法律判断，含真实人类数据对照，并指出偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":1,"question":"大语言模型聊天机器人在法律合理性判断上是否与人类判断一致？","design":"比较26个LLM（来自Meta、Google、Anthropic、OpenAI、DeepSeek、xAI）与人类被试在25个法律相关合理性场景中的回答分布，分析模型回答的集中度、对政府与企业的偏向性，以及与不同人口群体回答的一致性。","baseline":"人类被试对相同25个法律合理性场景的回答数据。","findings":"LLM的回答分布与人类有统计差异，但总体落在人类回答的经验范围内，大体追踪人类的合理性观念。LLM回答更同质化，有时将可变标准当作不变规则，且更偏向政府和企业，并与白人、男性、年长、高教育程度人群的回答更一致。","reliability":"论文指出需要更系统的研究来确认或否定初步发现，并承认LLM回答存在同质化、偏向性等潜在问题。","relevance":"该研究直接比较LLM与人类法律判断，包含真实人类数据对照，并揭示了仿真偏差，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其多模型比较和人口统计学对齐分析的方法，评估LLM在特定判断任务中的偏差。｜可迁移到信贷审批歧视、消费者投诉处理或监管合规判断等经济金融场景。｜以LLM作为虚拟信贷员，处理贷款申请并给出批准决策，结果变量为批准率及理由，与真实银行信贷审批数据对照，检验模型是否与人类决策一致及是否存在人口统计学偏差。"}},{"id":"2603.17094","version":2,"title":"Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction","zh_title":"评估LLM模拟对话在建模人类社交互动中不一致与不合作行为的表现","abstract":"Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation of simulated conversations by explicitly recognizing that human conversations inherently involve inconsistent and uncollaborative behaviors, such as misunderstandings and interruptions. Since these behaviors contribute to the complexity of human social interaction, we argue that LLM-simulated conversations should reproduce them at frequencies comparable to those observed in human conversations. To support a detailed and interpretable evaluation of these behaviors, we introduce CoCoEval, a framework consisting of an evaluation scheme based on turn-level detection of 10 types of inconsistent and uncollaborative behaviors and a benchmark for simulating conversations in professional scenarios involving collaboration and conflict. Using CoCoEval, we compare human conversations with those simulated by GPT-4.1, GPT-5.1, and Claude Opus 4. The results show that (1) LLM-simulated conversations exhibit far fewer inconsistent and uncollaborative behaviors than human conversations under vanilla prompting, and (2) prompt engineering and supervised fine-tuning do not provide reliable control over these behaviors, often leading to the overproduction of specific behaviors. CoCoEval identifies gaps between human and LLM-simulated conversations that are not captured by conventional evaluation based on conversation-level Likert scales, raising concerns about the use of LLMs as proxies for human social interaction.","authors":["Ryo Kamoi","Ameya Godbole","Binglin Zhou","Xiaoxin Lu","Longqi Yang","Rui Zhang","Mengting Wan","Pei Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-03-17","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2603.17094","pdf_url":"https://arxiv.org/pdf/2603.17094","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","对话模拟","算法保真度"],"reason":"直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":2,"question":"如何评估LLM模拟对话中不一致和不合作行为的保真度，并与人类对话对照？","design":"使用GPT-4.1、GPT-5.1和Claude Opus 4在专业协作与冲突场景下模拟30轮对话延续，通过普通提示、分类引导提示和监督微调三种设置，测量10类不一致和不合作行为的出现频率。","baseline":"来自QMSum、NCPC、SIM和IQ2数据集的真实人类对话，涵盖商业、学术、政府会议和辩论。","findings":"普通提示下LLM模拟对话的不一致和不合作行为远少于人类对话；提示工程和监督微调无法可靠控制这些行为，常导致特定行为过度产生。","reliability":"论文指出LLM模拟对话的行为频率高度依赖模拟设置，且所有评估设置均未能复现人类对话中这些行为的频率；传统对话级Likert量表无法捕捉这些差异。","relevance":"该研究直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"借鉴其细粒度行为检测评估方案和多种提示/微调设置对比，可迁移到经济金融中的协商、谈判或政策沟通模拟场景。｜例如，模拟消费者与客服的讨价还价或投资者与理财顾问的风险沟通。｜以LLM模拟谈判双方，处理为不同提示策略（如普通提示 vs. 明确要求包含冲突行为），结果变量为不一致行为（如误解、打断）的频率，并与真实谈判对话语料（如法庭记录或客服录音）对照。"}},{"id":"2607.27232","version":2,"title":"Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups","zh_title":"同情框架：跨社会人口群体评估AI对齐","abstract":"Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.","authors":["Haran Shani-Narkiss","Michael Fire","Oren Tsur"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-07-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2607.27232","pdf_url":"https://arxiv.org/pdf/2607.27232","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","情绪感知"],"reason":"用LLM复现人类情绪感知，并与大规模人类调查对照，评估对齐差异","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":2,"question":"大语言模型在多大程度上与人类对新闻标题中同情性框架的情感感知保持一致，这种一致性在不同社会人口群体间是否存在差异？","design":"以七个主流大语言模型（GPT-5.2、Grok、GPT-4、Gemini、DeepSeek、Claude、Mistral）作为“读者”，对216条涉及政治和地缘政治冲突的新闻标题进行二元判断（是否对冲突中某一方产生同情），并与3011名英国成年人的调查回答进行对比，测量模型与人类判断的斯皮尔曼相关性。","baseline":"通过YouGov对3011名具有英国人口代表性的成年人进行调查，收集了超过20万条人类对新闻标题同情性框架的二元评价。","findings":"模型与人类判断的一致性因模型而异，GPT-5.2最高（0.789），Mistral最低（0.41）；即使总体一致性很高，在年龄、教育、母语、政治意识、话题知识和既有观点等子群体间仍存在显著差异。","reliability":"论文承认总体对齐高并不代表对所有人口群体都一致，顶尖模型在老年人、低教育水平、非英语母语者、低政治意识和无强烈观点者中一致性较低；同时指出模型在不同话题上的对齐稳定性不同，某些模型存在特定领域失效。","relevance":"该研究直接回应了用LLM替代人类被试进行情感感知仿真的可靠性问题，提供了大规模、人口多样化的真实人类对照，并揭示了总体对齐掩盖的群体差异，对评估LLM在调查和实验中的适用性具有重要参考价值。","inspiration":"借鉴其大规模分层对照设计，将模型输出与人口代表性样本的个体级回答进行相关性分析，并检验子群体差异，以识别对齐失效的边界条件。｜可迁移到政策公告的预期形成研究，例如央行沟通中文本框架对公众通胀预期的影响。｜以LLM作为虚拟受访者，对央行声明进行情感或框架判断，与家庭调查中的通胀预期数据（如密歇根大学消费者调查）对照，检验模型在不同教育、年龄和金融素养群体中的对齐程度。"}},{"id":"2609.11108","version":1,"title":"But How Would AI Agents Run a Town's Economy?","zh_title":"AI代理如何管理城镇经济？","abstract":"We placed 100 memory-equipped large language model (LLM) agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography (earning wages, running businesses, setting prices) and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies. Across 91 validated runs (2.44M agent decisions, 21.5B tokens), the money stops moving, in a specific and measurable way. A 12x tourist demand shock raises business revenue 4.62x ($p<0.001$), which we decompose exactly into a 1.50x extensive margin (more businesses trading) and a 3.07x intensive margin (more revenue each). Monetary transmission stops there. Wages move 1.03x ($p=0.42$); 0.3% of 3,981 menu items are ever repriced ($p=0.47$). A randomized cash transfer (NPR 5,000 to 20 of 100 agents) shows the same pattern from the opposite direction: 96.7% is still held 311 pulses later, marginal propensity to consume 3-4% by two independent measures, indistinguishable from zero. The wealth distribution is consequently near-frozen at the horizon this literature uses ($\\rho=0.964$ over 2 simulated weeks), but not frozen. $\\rho$ falls to 0.832 at 12 weeks and 0.752 at 26, a horizon-dependence no short study can see. Matched ablations show which knob actually matters. Swapping the backing LLM moves every outcome we measure ($p=0.0039$); deleting agents' memory moves none of them detectably. A purely social tool fails 94-97% of the time across two model families, compared with ~96% success on economic tools, with no measurable shift away from it. Every headline number is verified twice, by a live validator and by an offline recomputation that reconciles each agent's wealth against its own signed transaction history, and we release the full run corpus for reanalysis.","authors":["Sajal Regmi","Siddhartha Pudasaini","Chetan Phakami Pun"],"categories":["cs.MA","cs.ET"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11108","pdf_url":"https://arxiv.org/pdf/2609.11108","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B3","B4"],"tags":["LLM代理","经济仿真","政策评估"],"reason":"用LLM代理模拟经济，与真实数据对照，涉及政策评估，并批判性指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":4,"question":"LLM智能体在封闭货币经济中能否实现货币流通与财富再分配？","design":"100个带记忆的LLM智能体在真实博卡拉湖滨地理空间上运行封闭货币经济，从事工作、经营企业、定价交易；施加旅游需求冲击和随机现金转移两种处理；测量企业收入、工资、价格调整、边际消费倾向和财富分布变化。","baseline":"无对照","findings":"货币流通在企业和家庭层面均受阻：旅游冲击使企业收入增加4.62倍，但工资仅变动1.03倍，价格几乎不调整；随机现金转移的边际消费倾向仅3-4%，与零无显著差异。财富分布在2周内几乎冻结，但延长至26周后缓慢放松，表明短期研究可能得出误导性结论。","reliability":"论文承认种子数偏少（每条件9个），依赖匹配种子或组内对比；模型选择显著影响结果，而记忆删除无影响；社会工具失败率高达94-97%，经济工具成功率约96%；仅使用特定LLM模型，结果可能不具普遍性。","relevance":"该研究直接回应了LLM仿真在经济学中的可靠性问题，通过随机实验和长期追踪揭示了仿真失效的具体机制，对评估LLM作为人类被试替代品的有效性具有重要参考价值。","inspiration":"借鉴其随机现金转移和旅游需求冲击的准实验设计，以及通过长短期对比揭示时间尺度依赖性的方法｜可迁移到消费者跨期选择、政策刺激的乘数效应或信贷扩张的传导机制等宏观金融场景｜以LLM智能体为被试，施加一次性收入转移或信贷额度提升，测量消费支出、储蓄率和资产配置变化，并与家庭金融调查或信用卡交易数据对照。"}},{"id":"2609.11611","version":1,"title":"Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust","zh_title":"生成式AI进入交通领域时谁承担风险？算法公平、合成数据有效性与公众信任的分配式社会技术审计","abstract":"Generative artificial intelligence is entering transportation through traveler-facing advisories, synthetic crash-record generation, and policy decision support. Existing governance frameworks lack transport-specific statistical tools to measure distributional risks across heterogeneous populations. We develop a Distributional Sociotechnical Audit (DSA) that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline. The audit analyzes 5,760 persona-controlled queries to four LLM families across 12 demographic cues and four transport topics, uses two cross-family judges and a Wasserstein-2 Equity Dispersion Index, tests three FARS crash-record generators with conditional projected maximum mean discrepancy (cpMMD), fits a Bayesian ordered-logit model to Pew American Trends Panel Wave 152 (N = 4,538), and combines the signals into a continuous Sociotechnical Risk Index. Congestion-pricing advice has the highest persona-based dispersion (mean EDI = 1.96; highest direct EDI = 2.20). CART synthetic crash records fail all conditional tests (p < 0.001), while the Gaussian copula has borderline conditional stress (p = 0.105) despite passing marginal checks. Attitudes to AI vary across demographic strata. Distributional audits and continuous risk indices with sensitivity reporting offer a more defensible basis for transport GenAI governance than categorical approval tiers, which show a 75% assignment flip rate under weight perturbation.","authors":["Amir Rafe","Subasish Das"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11611","pdf_url":"https://arxiv.org/pdf/2609.11611","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","交通政策","公平性审计"],"reason":"用LLM模拟公众对交通政策的态度，并与真实调查数据对照，评估公平性和有效性。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":5,"question":"生成式AI进入交通领域后，其输出、数据产品和公众态度在不同人群间的分布性风险如何测量与治理？","design":"对四个LLM家族施加12种人口学线索和4个交通主题的5,760个查询，用两个跨家族裁判和Wasserstein-2公平离散指数测量建议分布差异；用条件投影最大均值差异检验三种FARS事故记录生成器的条件分布有效性；用贝叶斯有序logit模型拟合Pew调查数据分析公众AI态度异质性；最后合成连续的社会技术风险指数。","baseline":"Pew American Trends Panel Wave 152（N=4,538）的真实调查数据，以及FARS的110,001条真实事故记录。","findings":"拥堵收费建议在不同人格间的分布离散度最高（平均EDI=1.96），政策争议话题的人格驱动变异比天气安全建议高1.6倍；CART合成事故记录未通过所有条件检验，高斯copula在边际检验通过的情况下条件压力检验边缘显著（p=0.105）。","reliability":"论文承认分类审批层级在权重扰动下存在75%的分配翻转率，因此主张采用带敏感性报告的连续风险指数；但未详细讨论LLM仿真在何种条件下会失效，也未系统检验人格线索与真实人群行为的一致性。","relevance":"该研究用LLM模拟不同人口学群体对交通政策的反应，并与真实调查数据对照，评估公平性和有效性，直接回应了研究者对LLM仿真可靠性及偏差的关注，值得精读其审计框架和统计方法。","inspiration":"值得借鉴的是用多组人格线索系统探测LLM输出分布差异，并用真实调查数据校准公众态度异质性，同时用分布距离指标而非简单准确率来度量公平性｜可迁移到政策公告的预期形成或消费者对金融产品的态度异质性研究，例如不同人口群体对通胀或利率变化的反应差异｜设计雏形：用LLM扮演不同收入、教育、年龄的消费者，施加不同措辞的央行政策声明，测量其通胀预期和消费意愿，并与密歇根消费者调查或纽约联储SCE的真实数据对照，检验LLM仿真能否复现真实人群的态度分布和异质性。"}},{"id":"2608.28482","version":2,"title":"How Proper Scoring Rules Shape LLM Forecasting","zh_title":"适当评分规则如何塑造大语言模型预测","abstract":"This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured.","authors":["Benjamin Turtel","Paul Wilczewski","Kris Skotheim","Ville A. Satop\\\"a\\\"a","Philip E. Tetlock"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-11","first_seen":"2026-08-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2608.28482","pdf_url":"https://arxiv.org/pdf/2608.28482","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["LLM预测","校准与偏差","评分规则"],"reason":"评估LLM预测校准与偏差，有真实事件结果对照，涉及统计推断有效性","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":3,"question":"不同的严格恰当评分规则作为训练奖励时，如何影响大语言模型预测者的性能与行为？","design":"使用GPT-OSS-120b模型，在二元真实世界事件预测任务上，分别以对数、Brier、球面、Beta(0.5,0.5)和Beta(2,2)五种严格恰当评分规则作为Dr. GRPO的终端奖励进行训练，比较各模型在聚合评分、校准、区分度、概率使用及偏差-信息-噪声分解上的差异。","baseline":"无对照","findings":"不同奖励训练的模型在聚合准确性和区分度上差异较小，但在校准、概率使用和偏差-信息-噪声构成上存在明显差异；Brier训练的模型Brier分数最低且AUC-ROC最高，而对数训练的模型对数分数最高且校准误差最低。","reliability":"论文未讨论","relevance":"该研究直接评估LLM预测的校准与偏差，使用真实事件结果作为基准，与您关注的经济学预测和政策评估场景高度相关，值得精读原文以了解奖励设计如何影响仿真可靠性。","inspiration":"借鉴其通过改变训练奖励函数来塑造模型预测行为的方法，可系统比较不同激励下的预测偏差结构｜可迁移到经济预测场景，如通胀预期、政策效果或资产价格预测，考察不同损失函数对预测者行为的影响｜以LLM作为经济预测者，分别用对数、Brier等评分规则微调，比较其在CPI预测或政策公告效应预测中的校准与偏差，并与专业预测者调查数据（如SPF）对照。"}},{"id":"2609.11198","version":1,"title":"(Whose defaults?) Is artificial intelligence reorienting archaeological methods?","zh_title":"谁的默认？人工智能正在重新定向考古学方法吗？","abstract":"Generative AI and the practice of \"vibe coding\" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?","authors":["Lorenzo Cardarelli","Roberto Ragno"],"categories":["cs.CY","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11198","pdf_url":"https://arxiv.org/pdf/2609.11198","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","方法多样性","科学实践"],"reason":"评估LLM对考古方法选择的影响，有真实文献数据对照，并批判性指出收敛风险，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":5,"question":"生成式AI和“氛围编程”是否正在收窄考古学研究中计算方法的多样性？","design":"本研究并非用LLM模拟人类被试，而是评估LLM对方法选择的影响。第一部分分析2010-2025年约119,000篇考古学摘要，用本地LLM提取并聚类计算技术，再用贝叶斯狄利克雷-多项模型检验2023年后方法构成的变化。第二部分为受控实验：让两个开源LLM为28个标准化考古研究问题推荐方法，提示词分新手、中级、专家三个方法学指导水平，测量推荐多样性与文献的差异，并用负二项回归检验方法推荐频率是否受2023年前流行度影响。","baseline":"真实人类基准为Scopus中2010-2025年考古学文献摘要所反映的实际方法使用分布，以及2023年前的方法流行度数据。","findings":"文献分析显示2023年后方法构成有微小但可信的偏移，但远小于领域内既有异质性，且整体方法多样性上升而非下降。LLM推荐的方法多样性远低于文献，尤其在新手提示下更低，且更偏向2023年前已流行的方法，其推荐模式更接近2023年后的文献。","reliability":"论文承认研究设计无法建立因果关系，只能说明结果与LLM导致方法收敛的压力一致；LLM推荐实验使用标准化问题，可能未完全反映真实研究情境的复杂性；且仅使用两个开源模型，结果可能不具普遍性。","relevance":"该研究虽非直接的人类仿真实验，但通过真实文献数据对照和受控实验，批判性地揭示了LLM对方法选择的收敛性影响，对关注LLM在学术研究中偏差与可靠性问题的研究者有参考价值。","inspiration":"借鉴其受控实验设计：通过不同提示词水平（新手/专家）施加处理，测量LLM输出的多样性并与真实数据分布对比，可迁移到经济金融研究中LLM辅助方法选择或模型设定的场景。｜例如，在资产定价研究中，可让LLM为不同经验水平的研究者推荐计量模型或因子选择，检验其是否导致方法同质化。｜设计：以经济学博士生为被试，随机分配使用LLM辅助或传统方法进行实证分析，处理为是否提供LLM推荐，结果变量为所选计量方法的多样性和与文献分布的偏差，对照真实数据为顶级经济学期刊中实际使用的方法分布。"}},{"id":"2609.10939","version":1,"title":"Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training","zh_title":"评估面向脚手架的多智能体大语言模型系统用于临床访谈训练","abstract":"Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.","authors":["Luming Yang","Haoxian Liu","Siqing Li","Rong Jia","Yue Xiao","Guanhua Chen","Li Lu"],"categories":["cs.MA","cs.AI","cs.HC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10939","pdf_url":"https://arxiv.org/pdf/2609.10939","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医学教育","多智能体"],"reason":"用LLM模拟标准化病人训练医学生，有真实学生对照，属人类仿真但场景为教育训练而…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":4,"question":"多智能体LLM系统能否通过脚手架式提示提升医学生的临床访谈过程质量，而不影响诊断准确性？","design":"用三个专用LLM智能体（患者、导师、逐轮评估者）模拟标准化病人并提供苏格拉底式提示，对100名医学生进行随机对照实验，比较脚手架条件与结构化非LLM控制条件，结果用OSCE对齐的评分量表测量。","baseline":"对照组为接受结构化渐进信息呈现的非LLM训练条件的医学生，最终考试在仅患者环境中进行，以OSCE评分作为真实人类表现基准。","findings":"多智能体AI标准化病人系统显著提高了最终考试中的沟通、共情表达和特定病史采集行为得分，但诊断准确性与对照组无显著差异。","reliability":"论文未讨论。","relevance":"该研究用LLM模拟人类被试（标准化病人）并设置真实人类对照（医学生），属于人类仿真在教育训练场景的应用，对关注LLM仿真可靠性与偏差的研究者有参考价值，但非经济学或政策评估场景。","inspiration":"借鉴其多智能体分工和脚手架式干预设计，将LLM模拟角色与评估角色分离，以过程质量而非最终结果作为主要结果变量。｜可迁移到经济金融中的沟通与决策训练场景，如信贷审批中的客户沟通、金融咨询中的信息披露与信任建立。｜以LLM模拟客户或投资者，对经济学专业学生或从业者进行随机分组，处理组接受多智能体LLM的苏格拉底式提示与逐轮反馈，对照组接受传统案例学习，结果变量为沟通质量、信息获取完整性和决策准确性，并与真实客户互动数据或专家评分进行对照。"}},{"id":"2609.10856","version":1,"title":"Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation","zh_title":"遵循偏好，错失最优：AI住房推荐中的合规无优化","abstract":"Large language models are becoming the first point of contact for consumer search in domains where the stakes are material and the law is explicit. Existing audits show that models steer housing seekers by perceived identity, but none can say what a user loses when a recommender overlooks a suitable option, for want of an enumerated inventory to score omissions against. We audit AI housing recommendation against a verifiable ground truth. For each of 150 synthetic renter scenarios in New York City we build a pool of 120 real listings with known rent, bedrooms and GTFS-computed transit commute, compute the exact set satisfying the renter's stated constraints, and derive its Pareto frontier. The primary outcome assumes no utility function: a recommendation is strictly dominated if the same pool holds a listing cheaper, faster to commute from and no smaller in bedrooms. Across 9,945 calls to three models from two vendors, compliance is near-perfect (1.8% violation against a 66.6% random floor), yet 39.0% of recommendations are strictly dominated, and the dominating listing is a median 900 USD/month cheaper and 3.5 minutes closer. A within-scenario manipulation separates two capabilities usually conflated: changing one sentence moves median recommended rent by 646 USD/month in the correct direction, so preferences are honored, yet recommendations still sit 606 USD/month above the five cheapest qualifying listings on the same screen, and an unambiguous lexicographic instruction gives no improvement under equivalence testing against a pre-specified 50 USD/month bound. The gap widens with candidate-set size and replicates across OpenAI and Anthropic models to within 3 USD. We characterize the failure as compliance without optimization, propose dominance-rate instrumentation as a deployable diagnostic, and release all code, prompts and per-call results.","authors":["Hsuan Lo"],"categories":["cs.CY","cs.IR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10856","pdf_url":"https://arxiv.org/pdf/2609.10856","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","推荐系统审计","决策偏差"],"reason":"用LLM模拟住房推荐中的用户决策，与真实房源数据对照，评估合规性与优化缺失，可…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":10,"question":"在住房推荐场景中，大语言模型是否在遵守用户明确约束的同时，未能优化推荐结果，导致用户错失更优选项？","design":"该研究并非用LLM模拟人类被试，而是直接审计LLM作为住房推荐系统的行为。研究者构建了150个纽约市租房者场景，每个场景配120个真实房源，计算满足约束的集合和帕累托前沿。然后调用三个模型（OpenAI和Anthropic）共9945次，要求模型推荐房源，并测量推荐是否违反硬约束、是否被严格占优（存在更便宜、通勤更快且卧室数不更少的房源）、以及租金与最优的差距。还通过改变场景中的偏好句子来检验模型是否遵循偏好，并通过改变候选集大小来检验优化能力。","baseline":"无对照","findings":"模型几乎完美遵守硬约束（违规率1.8%），但39.0%的推荐被严格占优，占优房源中位数便宜900美元/月且通勤快3.5分钟。模型遵循偏好（改变一句话租金中位数变动646美元），但推荐仍比同屏最便宜五个合格房源贵606美元，且明确的词典序指令无改善。","reliability":"论文承认其身份条件差异的零结果仅适用于固定候选集重排，不适用于开放式搜索；且未讨论模型在真实用户交互中的表现或长期影响。","relevance":"该研究直接评估LLM在住房推荐中的决策质量，与人类真实房源数据对照，揭示了合规与优化的分离，对关注LLM仿真可靠性及偏差的研究者有重要参考价值。","inspiration":"借鉴其构建可验证真实数据集并计算帕累托前沿来量化机会成本的方法，以及通过改变提示中的偏好来分离合规与优化能力的设计。｜可迁移到信贷审批或保险定价场景，检验LLM是否在遵守申请人硬性条件的同时未能推荐最优贷款或保单。｜用LLM扮演信贷员，输入申请人特征和贷款产品池，要求推荐产品；结果变量为推荐产品是否被占优（存在利率更低、费用更少且额度不低的产品），对照真实贷款产品数据和申请人约束，测量占优率和成本差距。"}},{"id":"2608.21057","version":2,"title":"Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment","zh_title":"通过人类对齐设计用于药物发现中智能体 AI 的稳健 LLM 评估系统","abstract":"Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with five contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate the judge through a human alignment study with five expert annotators, comparing Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as candidate judges. Third, we optimize the best-performing judge using few-shot demonstrations of human-annotated examples, improving alignment with the human majority vote from 0.80 to 0.86. Fourth, applying the optimized judge to 70 held-out questions, we surface concrete limitations and find no strong evidence that informal phrasing degrades output quality; it may, however, still be helpful to have the LLM rewrite the original question before querying the agent. Finally, we extend the framework to 38 adversarial questions that are ambiguous, invalid, out-of-scope or ethically sensitive, and show that the agent's refusal behavior is guided by the stated intent of a request. Our framework provides a reusable template for human-aligned evaluation of agentic systems in scientific domains.","authors":["Emma Granqvist","Roc\\'io Mercado","Samuel Genheden"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-11","first_seen":"2026-08-24","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2608.21057","pdf_url":"https://arxiv.org/pdf/2608.21057","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM-as-a-Judge","人类对齐","药物发现"],"reason":"LLM 作为评判者替代人工评估，属于标注替代而非仿真人类被试，但涉及人类对齐验…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2609.02990","version":2,"title":"Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy","zh_title":"面向参与式民主的偏好推断集体中心评估","abstract":"To scale up collective decision-making, participatory democracy platforms such as Polis and Remesh enable online deliberation among thousands of participants. However, at this scale, participants cannot review every opinion submitted by others, producing highly sparse voting data that misrepresent patterns of consensus, conflict, and minority support. Platforms therefore increasingly rely on Preference Inference (PI) models to predict missing votes. Yet this automation is not neutral: inferred preferences can artificially amplify, suppress, or reorder existing patterns of support, ultimately reshaping how the outcomes of a deliberation are interpreted. More generally, we lack a systematic understanding of how existing PI methods affect the collective preference landscape. To address this gap, we benchmark several existing PI approaches in this context. Moving beyond conventional user-centric evaluations centered on the accuracy of individual predictions, we introduce a collective-centric evaluation framework that measures whether inferred votes preserve salient properties of the broader preference landscape. We further contribute the largest multilingual dataset of its kind: four consultations spanning over 90k participants, 1M votes, and 22 languages. Our experiments show that models with comparable predictive accuracy can differ substantially in the degree to which they preserve the collective structure. These results demonstrate that accuracy alone is insufficient for evaluating PI in democratic settings. By contributing a novel comprehensive and collective-centric evaluation benchmark for the task of PI, this work aims to support the development of AI systems that scale deliberation without compromising the integrity of its democratic outcomes.","authors":["Pierre-Antoine Lequeu","Salim Hafid","Paul Lerner","Nazanin Shafiabadi","Laur\\`ene Cave","David Mas","Jean-Philippe Cointet","Benjamin Piwowarski","Fran\\c{c}ois Yvon"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-04","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.02990","pdf_url":"https://arxiv.org/pdf/2609.02990","source_feed":"cs.SI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["偏好推断","参与式民主","集体决策"],"reason":"涉及社会模拟但无LLM仿真人类被试，且无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2609.11067","version":1,"title":"When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text","zh_title":"当噪声制造偏见：噪声文本下LLM作为评判者的偏见测量脆弱性","abstract":"Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.","authors":["DongHyun Ryu","Jaehyeok Lee","YeongJun Hwang","JinYeong Bak"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11067","pdf_url":"https://arxiv.org/pdf/2609.11067","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见测量","噪声鲁棒性","算法公平性"],"reason":"研究LLM作为评判者测量文本偏见，属于对LLM本身测量属性的评估，而非用LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.10883","version":1,"title":"Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble","zh_title":"故事印记：AI助手从相似的人类角色中吸收特质","abstract":"Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's behavior in multi-turn conversations with users, a format quite different from the stories? And does the Assistant adopt the behaviors and preferences of human characters? We refer to this adoption as story imprinting. We finetune GPT-4.1 and Kimi-K2.6 on stories in which generally helpful human characters give subtly harmful advice after being insulted. The Assistant adopts the same conditional behavior while otherwise remaining helpful. This occurs even when fewer than 2% of stories depict the behavior. In a separate experiment, the Assistant adopts preferences that are only implicit in the narration. A human character's body language suggests they dislike working on spreadsheets, yet they never say so and continue giving good advice on spreadsheets. After finetuning, the Assistant becomes less likely to choose spreadsheet tasks. Next we ask which characters most influence the Assistant. We find the Assistant adopts behaviors more often from characters that resemble it (e.g., helpful rather than dismissive). We call this the affinity effect. The effect extends to other personas elicited with system prompts: unhelpful personas adopt behaviors from unhelpful characters. We also observe it in finetuned base models. We use the affinity effect to learn how models represent the Assistant. We find the Assistant adopts behaviors more from characters affiliated with elite universities (e.g., Yale) than non-elite ones. This implies the model's internal representation of the Assistant is more similar to humans from elite universities. Overall, the Assistant can be influenced by stories that depict only human characters (no AIs), which may conflict with the Persona Selection Model for the Assistant.","authors":["Jorio Cocola","Lev McKinney","Harry Mayne","Jan Betley","Owain Evans"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10883","pdf_url":"https://arxiv.org/pdf/2609.10883","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["模型行为","人格测量","故事微调"],"reason":"研究AI助手从故事中吸收人类特质，属于模型行为测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.11109","version":1,"title":"How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding","zh_title":"AI编码员如何讨论、分歧并达成共识：基于LLM的定性编码的挑战与机遇","abstract":"The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.","authors":["Jeongyeon Kim","John Mitchell"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11109","pdf_url":"https://arxiv.org/pdf/2609.11109","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM定性编码","多智能体协作","标注替代"],"reason":"LLM替代人工编码员，属标注替代而非仿真被试，但涉及多智能体协作与人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.11529","version":1,"title":"Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration","zh_title":"伦理培训智能体：通过角色扮演与讨论促进基于群体的伦理教育，以实现伦理反思与探索","abstract":"Group-based ethics training for Science, Technology, Engineering and Mathematics (STEM) students is a complex challenge, requiring substantial resources and expertise. While activity-based teaching methods, such as role-playing and discussions, are commonly employed to simulate real-world scenarios, current practices are often manual and lack integration with effective online platforms for supporting group-based ethical discussions. In this work, we propose Ethics Training Agents, a group discussion system that leverages multiple LLM participants embodying distinct ethical orientations, along with a moderator agent, to enable structured human-AI group ethical discussions for collaborative reflection. We conduct a user study with 45 undergraduate STEM students to evaluate the learning outcomes and user experience. The results show that our system supports engagement, coordination, and perspective-taking in group discussions and has a positive influence on ethical sensitivity. We also discuss practical design strategies for integrating multiple LLM agents into multi-human group settings to facilitate ethics training for STEM students.","authors":["Youngseok Seo","Sueun Jang","Hyesoo Park","Renz Samuel Gutierrez","Joseph Seering","Uichin Lee"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11529","pdf_url":"https://arxiv.org/pdf/2609.11529","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM角色扮演","伦理教育","人机群体讨论"],"reason":"LLM扮演伦理角色与人类讨论，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.10724","version":1,"title":"Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge","zh_title":"完成任务还不够：在累积挑战下评估智能体的韧性与体贴参与","abstract":"Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.","authors":["Yuanchen Bai","Zijian Ding","Angelique Taylor"],"categories":["cs.AI","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10724","pdf_url":"https://arxiv.org/pdf/2609.10724","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","韧性评估"],"reason":"模拟医疗工作流中agent行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:09","error":null,"has_summary":false,"summary":null},{"id":"2606.20041","version":2,"title":"AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models","zh_title":"AI经济学家智能体：基于RAG、知识图谱和大语言模型的循证经济金融分析框架","abstract":"We propose an AI economist agent for economic and financial scenario analysis. Scenario design often requires analysts to assess emerging risks with limited historical precedent, combine information from many sources, and translate qualitative mechanisms into internally consistent quantitative paths. Large language models (LLMs) can search and synthesize this information, but fluent narratives alone do not establish the model-based calculations needed for economic conclusions. Our framework uses LLM agents to plan the analysis, retrieve relevant evidence, and organize economic mechanisms, while registered quantitative models generate numerical outcomes and predefined tests determine whether intermediate results can be used in the final report. We apply the framework to European macro-financial stress scenarios and bank capital analysis. The empirical analysis evaluates retrieval of economic mechanisms, scenario construction, model execution, and report generation under a historical information cutoff. The results show how the AI economist agent can combine flexible evidence retrieval and scenario construction while keeping the resulting analysis linked to identifiable sources and explicit model calculations.","authors":["Masahiro Kato"],"categories":["econ.GN","cs.AI","cs.LG","q-fin.EC","q-fin.GN"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-11","first_seen":"2026-06-18","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2606.20041","pdf_url":"https://arxiv.org/pdf/2606.20041","source_feed":"cs.LG","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","经济分析","RAG"],"reason":"多智能体协作完成经济分析任务，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11101","version":1,"title":"ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation","zh_title":"ProMediConv：法律纠纷调解中主动对话智能体的基准测试","abstract":"Dispute mediation is essential for maintaining social harmony and resilience, yet developing skilled mediators is costly and time-consuming. Existing LLM-based mediation research remains limited by unrealistic task formulations, low-fidelity datasets, and coarse evaluation metrics that obscure turn-by-turn dynamics. To address these gaps, we introduce ProMediConv, a novel benchmarking framework that models mediation as a proactive, multi-stage, and party-aware dialogue process incorporating 11 mediation strategies and four party behavior pattern (BP) states. Using 972 complete real-world cases, we construct a high-fidelity mediation dataset with utterance-level annotations of strategies and BP states. Furthermore, to better assess agent impact, we propose MAD (Mean Attribute Difference), a fine-grained metric that captures BP shifts throughout the dialogue. Leveraging this framework, we establish a comprehensive benchmark by evaluating diverse models alongside our tailored baseline ProMediAgent. Extensive empirical analyses reveal critical behavioral phenomena and underscore the persistent challenges current models face in dynamic, multi-party mediation. Ultimately, ProMediConv provides a rigorous foundation and a vital quantitative standard for advancing AI-assisted conflict resolution. Our dataset and codebase are accessible at https://github.com/ZsWei66/ProMediConv_repo.","authors":["Zesheng Wei","Mengfan Li","Wenhao Liu","Yixin Zhang","Zilei Wang","Yang Deng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11101","pdf_url":"https://arxiv.org/pdf/2609.11101","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["对话智能体","法律调解","基准测试"],"reason":"构建调解对话基准，评估LLM对话能力，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.08585","version":2,"title":"Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations","zh_title":"自动化可模拟性的局限：LLM模拟器可以绕过解释","abstract":"Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poch\\'e et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.","authors":["Antonin Poch\\'e","Fanny Jourdan","Nils Feldhus","Qianli Wang","Jing Yang","Simon Ostermann","Nicholas Asher","Philippe Muller","Vera Schmitt"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.08585","pdf_url":"https://arxiv.org/pdf/2609.08585","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","LLM模拟器","NLP评测"],"reason":"评估LLM模拟器预测任务模型输出的能力，属于NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":28,"question":"自动化可模拟性评估中，LLM模拟器是否依赖捷径（如任务先验或标签泄漏）而非解释内容来预测任务模型输出？","design":"复现并扩展ConSim协议：用LLM模拟器（如Llama-3.1-8B、Qwen2.5-7B等）扮演解释接收者，输入任务模型的解释（概念、归因、理由等）和输入实例，要求预测任务模型的输出；在非匿名和匿名类名两种设置下，比较不同解释方法的可模拟性得分，并引入classes-as-concepts基线。","baseline":"无对照（未使用真实人类数据，仅与ConSim的LLM模拟结果进行定性比较）","findings":"在非匿名设置下，解释对模拟器预测的影响很小，模拟器主要依赖任务先验直接解决分类任务；在匿名设置下，classes-as-concepts基线（仅泄漏标签映射）表现优于所有测试的解释方法，表明匿名化可能奖励标签泄漏而非有意义的解释。","reliability":"论文承认其发现基于15B参数以下、无推理能力的LLM模拟器，可能不适用于更大或闭源模型；且实验任务可能已被训练数据污染，未处理污染问题。","relevance":"该研究直接评估LLM模拟器的可靠性，揭示其捷径行为，对使用LLM进行人类仿真实验的研究者具有重要警示意义，值得阅读原文以了解具体失效模式和稳健性建议。","inspiration":"借鉴其通过引入基线（如classes-as-concepts）和对比非匿名/匿名设置来检测捷径行为的方法，可迁移到经济金融领域的LLM仿真实验，如政策公告解读或信贷审批解释的仿真；设计雏形：用LLM模拟投资者或贷款申请人，处理为提供不同解释（如模型决策理由），结果变量为预测模型决策的准确率，对照真实人类实验数据（如调查或行为实验）以评估仿真有效性。"}},{"id":"2609.10758","version":1,"title":"Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu","zh_title":"徒有其名的多语言？乌尔都语中LLM的文化与语言弱点","abstract":"Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.","authors":["Farah Adeeba","Abdul Rafae Khan","Rajesh Bhatt","Hassan Sajjad"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10758","pdf_url":"https://arxiv.org/pdf/2609.10758","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["低资源语言","故事生成","模型评估"],"reason":"评估LLM生成故事的语言质量，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.10996","version":1,"title":"Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models","zh_title":"重新思考LLM作为评判者的口头置信度：2025年后专有模型的兼容性转变","abstract":"Verbalized confidence, long dismissed as overconfident, coarse, and prone to round-number clustering, is now the more robust soft-scoring mechanism for LLM-as-a-Judge on top-tier proprietary models. Across SummEval, AggreFact, and HelpSteer2, spanning up to 18 LLMs, we show that the standard advice to prefer log-probabilities no longer holds on post-2025 models, where verbalized confidence is the better signal. We call this a compatibility shift. On top of a standard verbalized-confidence baseline, we introduce two new ingredients: an overconfidence advisory and self-debate. Together they improve calibration, score-distribution spread, and robustness to task subjectivity. We further observe a generation effect: post-2025 models accommodate these two additions with little balanced-accuracy cost, whereas pre-2025 models pay a measurable penalty. Compared with logprob-based G-Eval, verbalized confidence is the more subjectivity-robust soft signal on GPT-family top-tier releases. The shift is invisible under accuracy-only reporting. Rather than defaulting to hard predictions, we recommend broader use of soft scoring in LLM-as-a-Judge. More broadly, verbalized confidence has moved from a weaker substitute for logprobs to a practical soft-scoring mechanism for contemporary LLM judges.","authors":["Yu-Chung Hsiao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10996","pdf_url":"https://arxiv.org/pdf/2609.10996","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判","置信度校准","NLP评测"],"reason":"研究LLM作为评判者的置信度校准，属于NLP评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.11020","version":1,"title":"K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models","zh_title":"K/V缓存干预分离解码器语言模型中的表征对齐与人格表达","abstract":"We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention configurations applied to Llama-3.1-8B for a fixed source-to-target persona pair, we report two consistent dissociations between representation-level alignment and behavioral expression, plus a common failure under position perturbations. First, all layer-band K/V replacements (early, mid, late) achieve strong local V-space alignment (V-gap 0.91, 0.89, 0.84), but only mid-layer replacement (layers 9-20) combines substantial target-marker expression with preserved lexical diversity. Second, full and mid-layer replacement induce comparable alignment (V-gap 0.94 vs. 0.89) yet produce different lexical-diversity profiles (TTR 0.65 vs. 0.77). Third, position perturbations (lag and shuffle) apply distinct operations yet uniformly suppress target-persona expression -- a common behavioral failure rather than a strict dissociation. Representation-level similarity metrics alone are thus not sufficient predictors of downstream persona expression in the regimes we study; the K/V cache emerges as a controllable but structurally constrained intervention surface. Because the transplanted trajectory carries the target's own generated token history, we characterize the intervention as trajectory-level transplantation rather than isolated persona-representation injection; a same-token-sequence control, decoding an identical token sequence under source vs. target conditioning, reproduces the sign and layer localization of the L28 representational shift, indicating the shift is not explained solely by imported token history. These findings characterize representation-behavior dissociation in a high-signal setting rather than establishing universality across models or persona pairs.","authors":["Yu Sun","Mengyin Lu","Cong Feng","Guangming Lu","Huimin Han"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11020","pdf_url":"https://arxiv.org/pdf/2609.11020","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","表征对齐","缓存干预"],"reason":"研究角色扮演中的人格表达控制，无实验或测量目的，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.11117","version":1,"title":"Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers","zh_title":"NLPCC 2026共享任务11概述：基于智能体的科学论文实验复现","abstract":"Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus on final repositories and are typically limited to machine learning (ML). We introduce AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains. Our framework uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics. AgentActionBench contains 150 papers, including 120 ML papers and 30 AI4Science papers. A human-annotated subset covering 10% of the benchmark provides validation data, while model-assisted augmentation expands the full benchmark to more than 10,000 rubric items. Experimental results show that current systems remain limited, with execution as the primary bottleneck. Meanwhile, the strong Pearson and Spearman correlations between model-generated and human-annotated rubrics validate the reliability of our scalable rubric-generation approach.","authors":["Hanhua Hong","Yizhi Li","Luu Gia Huy","Jian Yang","Ming Zhou","Chenghua Lin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11117","pdf_url":"https://arxiv.org/pdf/2609.11117","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体实验复现","基准测试","多智能体协作"],"reason":"该工作评估LLM智能体复现科学实验的能力，属于多智能体协作完成任务，不涉及以人…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.11769","version":1,"title":"Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing","zh_title":"识别不等于反转：事实保持型新闻框架的受控反转测试","abstract":"Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a model can undo a known framing transformation while keeping the facts fixed. We introduce a controlled inversion test over three established textual realizations of framing: evaluative lexis, agency realization, and information salience. Across 60 news articles and three intervention strengths, this yields 540 paired variants with preserved atomic facts and recorded edits. Across Qwen, DeepSeek, and Kimi, factual preservation remains near 0.84, whereas intervention reversal is 0.044--0.068. Even when both framing type and direction are recognized correctly, pooled reversal reaches 0.071. These results reveal a clear separation between factual fidelity, framing recognition, and framing inversion: recognizing how an article is framed does not imply that the framing can be undone.","authors":["Yi Liu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11769","pdf_url":"https://arxiv.org/pdf/2609.11769","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["新闻框架","LLM能力评测","文本改写"],"reason":"研究LLM改写新闻的框架反转能力，属NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11063","version":1,"title":"The information geometry of large language models is shared, learned, and controllable","zh_title":"大语言模型的信息几何是共享的、可学习的和可控的","abstract":"Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.","authors":["Dario Picozzi"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11063","pdf_url":"https://arxiv.org/pdf/2609.11063","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["信息几何","模型行为控制","表征分析"],"reason":"研究LLM输出几何与行为控制，不涉及人类被试仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.11900","version":1,"title":"MindTopo: Can Foundation Models Reason in Topological Space?","zh_title":"MindTopo：基础模型能在拓扑空间中推理吗？","abstract":"Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity, separation, order, enclosure, and knots. MindTopo evaluates each property at two cognitive levels. Reasoning asks a model to identify topological relations or infer how they change. Planning instantiates a foundation model as a closed-loop agent whose policy selects environment actions. MindTopo contains 11,030 instances across 13 procedurally generated task types with controllable difficulty. We benchmark 14 MLLMs and study agent configurations augmented with image and video generation, including 3 video generative models in planning settings. Every MLLM performs better on reasoning than on planning, and the best-performing model remains far below observed human performance. On Qwen3-VL-2B-Instruct, supervised fine-tuning and reinforcement learning improve reasoning more than planning. Generated observations retain local cues and reach plausible endpoints, but audited rollouts do not reliably follow environment dynamics or preserve topology across transitions. Our website is at https://mind-topo.github.io/","authors":["Yunfei Ge","Anbang Liu","Qineng Wang","Johnalbert Garnica","Jianwen Lyu","Zihan Wang","Reuben Tan","Jianfeng Gao","Ruohan Zhang","Yining Hong","Jiajun Wu","Manling Li"],"categories":["cs.AI","cs.CL","cs.CV"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11900","pdf_url":"https://arxiv.org/pdf/2609.11900","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["拓扑推理","多模态基准","认知评测"],"reason":"评估基础模型在拓扑空间推理与规划能力，属认知能力评测，非用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2609.11018","version":1,"title":"Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks","zh_title":"定义AI智能体：标准、指标与基准汇编","abstract":"The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, we examine how the underlying capability has been conceptualized across prior work and synthesize the metrics, benchmarks, and evaluation frameworks used to assess it. This review provides a structured account of the current landscape of agent evaluation, highlighting both established approaches and areas where evaluation remains limited or inconsistent. We additionally introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods identified through this review. Together, the survey and compendium provide a common structure for evaluating and comparing agent capabilities across AI systems, supporting more reproducible research, clearer communication, and more systematic study of artificial agents.","authors":["Mia Lassiter","Brinnae Bent"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11018","pdf_url":"https://arxiv.org/pdf/2609.11018","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI智能体","评估基准","综述"],"reason":"论文综述AI agent评估，不涉及用LLM仿真人类被试或与人类数据对照，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.11770","version":1,"title":"The widening evaluation gap in medical large language model research 2023 to 2026","zh_title":"医学大语言模型研究中的评估差距扩大：2023至2026","abstract":"Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.","authors":["Raad Bin Tareaf","Murad Al-Rajab","Samia Loucif"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11770","pdf_url":"https://arxiv.org/pdf/2609.11770","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学LLM","评估差距","研究设计"],"reason":"研究医学LLM评估差距，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11022","version":1,"title":"New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models","zh_title":"新证据，相同选择：测试视觉语言模型中的物理实验选择","abstract":"A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform. Current physical reasoning benchmarks usually evaluate only the final answer, so they do not directly measure this decision-making ability. We introduce a controlled evaluation where each problem provides one measurement image and four possible physical worlds created by combining two possible masses and two possible values of another relevant property. The model must either stop and answer or select the cheapest additional experiment that can resolve the question. We construct matched problem pairs where changing either the observed measurement or the question changes the optimal action. Since all possible worlds and experiment costs are known, we can explicitly determine the optimal choice. Across six open models and 144 physical parameter sets, direct responses repeat the same action for 95.1% to 100% of image pairs even when the correct action changes. Brief reasoning improves action switching, but the best model makes both decisions correctly for only 5.9% of image pairs. Additional analysis reveals failures in measurement interpretation, physical reasoning, and response formatting. By evaluating evidence selection separately from final answers, our benchmark reveals limitations in physical reasoning that conventional answer accuracy can overlook.","authors":["Sourajit Saha","Shubhashis Roy Dipta","Nobin Sarwar","Shaswati Saha","Yuxuan Jiang","Siyuan Li","Qiheng Wang"],"categories":["cs.CV","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11022","pdf_url":"https://arxiv.org/pdf/2609.11022","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","物理推理","基准测试"],"reason":"评估视觉语言模型在物理实验选择上的决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.11019","version":1,"title":"Work, Wellbeing, and Choice: Empirical Lessons for AI Futures","zh_title":"工作、幸福感与选择：AI未来的实证教训","abstract":"Advances in AI-driven automation have raised questions about how humans might find wellbeing in a world where paid employment is less necessary or less available than before. Paid work has been variously characterized as both a contributor and an impediment to human wellbeing. What is already known about the relationship between paid work and wellbeing? What factors influence wellbeing among people who do not work---or who do not need to work? And how might these factors bear upon prospective AI-induced economic transformations? To help provide empirical grounding for these questions, we survey the psychological, sociological, and economic literature that investigates the relationship between wellbeing and work. We draw on evidence from multiple populations, including the unemployed, retirees, lottery winners, and financially dependent spouses. This comparative review draws from studies across OECD countries, China, India, and Gulf states. We identify three key factors that mediate the relationship between work status and wellbeing: (1) agency and choice---whether the exit from work is voluntary or involuntary, as well as long-term agency; (2) the availability of alternative sources of work's latent benefits---such as volunteering, hobbies, or state-provisioned employment; and (3) social and systemic context---including cultural norms around work and the robustness of social safety nets. We draw on these three factors to derive specific implications for different AI automation scenarios, connecting the empirical evidence to concrete policy considerations.","authors":["Stephanie C. Y. Chan","Adam Bales","Katherine L. Hermann","Iason Gabriel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11019","pdf_url":"https://arxiv.org/pdf/2609.11019","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["工作与幸福感","AI自动化","文献综述"],"reason":"论文综述人类工作与幸福感关系，未使用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.11709","version":1,"title":"When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making","zh_title":"当智能体意见不一致时：贝叶斯反向推理作为多智能体集体决策的无标签锚点","abstract":"When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.","authors":["Ken Chen","Wei Wang","Sachith Seneviratne","Hansani Weeratunge","Saman Halgamuge"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11709","pdf_url":"https://arxiv.org/pdf/2609.11709","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","集体决策","贝叶斯推理"],"reason":"纯多智能体协作决策，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.09887","version":1,"title":"When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors","zh_title":"被告陈述何时重要？LLM模拟陪审员中的偏见与说服研究","abstract":"LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench","authors":["Cho-Ying Wu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09887","pdf_url":"https://arxiv.org/pdf/2609.09887","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","陪审团决策","法律偏见"],"reason":"用LLM模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策偏差与说服效应。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":1,"question":"在普通法陪审团审判中，被告的法庭陈述何时以及如何影响 LLM 模拟陪审员的裁决，重点关注说服、意识形态偏见和背景亲和力。","design":"使用 20 个前沿 LLM 模拟具有不同意识形态背景的陪审员，在 500 个争议性刑事案件中，对被告的不同背景和带有不同情感诉求或反驳的法庭陈述做出有罪/无罪及严重程度判断，并记录决策理由。","baseline":"无对照","findings":"情感说服可能适得其反，因为陪审员可能将其视为有罪或不一致的证据；背景契合度（陪审员与被告背景相似性）比孤立因素更强且显著，陪审员通常对背景相反的被告更严厉，对背景相同的被告更宽容；陪审员意识形态也强烈影响严重程度判断。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策中的偏差与说服效应，属于人类仿真实验，但未提供真实人类数据基准，可靠性存疑，值得阅读原文以了解其仿真设计细节和潜在偏差。","inspiration":"借鉴其通过系统操纵被告背景、陈述情感强度和陪审员意识形态来测量决策偏差的因子设计方法，以及大规模生成争议案例和记录决策理由的做法。｜可迁移到信贷审批歧视、招聘面试评估或政策沟通中的说服效应等经济金融场景。｜用 LLM 模拟信贷员或招聘经理，处理变量为申请人背景（如种族、性别）和陈述情感强度，结果变量为批准/拒绝或评分，对照真实信贷审批数据或审计研究结果来评估 LLM 仿真的外部有效性。"}},{"id":"2609.10280","version":1,"title":"Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models","zh_title":"总体模拟调查误差：设计和诊断大语言模型的调查回答","abstract":"Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theoretical and empirical case study, enables survey simulation designers to systematically identify and reflect on errors in LLM-generated surveys.","authors":["Indira Sen","Georg Ahnert","Leah von der Heyde","Jana Lasser","Bernd Wei{\\ss}","Markus Strohmaier"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10280","pdf_url":"https://arxiv.org/pdf/2609.10280","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","误差框架"],"reason":"提出TS2E框架诊断LLM调查仿真误差，含实证案例，直接相关","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":2,"question":"如何系统识别、追踪和记录大语言模型生成调查回答中的误差来源？","design":"提出一个概念框架（TS2E），将调查模拟生命周期划分为不同阶段，枚举各阶段可能出现的测量误差和代表性误差，并通过一个理论性和实证性案例研究进行说明。","baseline":"无对照","findings":"该框架区分了研究者设计选择导致的误差与LLM固有局限导致的误差，并引入了LLM特有的误差类型（如人物角色构建误差）和评估谬误。通过案例研究展示了框架如何帮助设计者系统识别和反思LLM生成调查中的误差。","reliability":"论文承认LLM训练数据存在缺口和偏斜，指令微调等后训练过程可能影响模型行为，且高总体对齐可能掩盖方差、子群体异质性和下游统计关系的严重失真。","relevance":"该论文直接针对LLM仿真调查的可靠性问题，提出了系统诊断误差的框架，对关注仿真效度与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其将总调查误差框架迁移到LLM仿真的思路，对仿真流程进行阶段分解并系统识别误差来源。｜可迁移到经济金融领域的调查类仿真，如消费者信心调查、通胀预期调查、投资者情绪调查等。｜设计雏形：用LLM模拟不同人口统计学特征的消费者，施加不同的经济信息提示（如货币政策公告），测量其通胀预期和消费意愿，并与密歇根大学消费者调查的真实数据对照，检验仿真误差。"}},{"id":"2609.09899","version":1,"title":"Strangers to Themselves: What Language Models Say About Themselves Is Generic","zh_title":"自我陌生：语言模型对自身的描述是泛化的","abstract":"Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about \"capable AI agents in general\" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.","authors":["Phil Blandfort","Urja Pawar"],"categories":["cs.LG","cs.AI","cs.CL","cs.CV","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09899","pdf_url":"https://arxiv.org/pdf/2609.09899","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM自我认知","行为预测","可靠性评估"],"reason":"评估LLM自我报告与行为的一致性，揭示自我认知偏差，对仿真可靠性有批判性启示。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":8,"question":"语言模型对自己行为的自我报告是否包含关于该模型自身的特权知识，还是仅仅反映了对AI助手的一般性理论加上有利偏差？","design":"该研究并非用LLM模拟人类被试，而是将LLM自身作为研究对象：在九项行为评估中测量模型在不同条件下的实际行为率，然后让模型预测这些行为率，并设置多种对照（如询问“一般AI智能体”、用其他模型的回答预测目标模型等），比较预测准确度。","baseline":"无对照（没有使用真实人类数据作为基准，而是以模型间共享结构和跨模型预测作为参照）。","findings":"模型的自报行为与实际行为相关性很弱（r=+0.04），即使提供具体评估题目也只提升到+0.24；而询问“一般AI智能体”的预测效果相当（+0.28），其他模型对目标模型的预测甚至更好。第一人称框架会系统性地使自报偏向有利方向，低估有害行为。","reliability":"论文指出，在模型间行为差异很小的评估上（如能力评估），自知识信号微弱，因此难以检测；模型的自报主要反映通用AI行为理论，而非模型特定知识；微调虽能改善特定行为的自预测，但会改变行为本身且不具泛化性。","relevance":"该研究对LLM仿真人类实验的可靠性有直接警示：若用LLM自我报告作为行为预测或态度测量的替代，可能仅得到通用模式而非个体特异性，且存在社会赞许性偏差。值得阅读原文以了解其对照设计。","inspiration":"借鉴其预测测试框架：将自我报告与行为测量分离，并设置通用主体、跨模型预测等对照，以剥离通用知识与自我知识。｜可迁移到经济金融中的个体偏好或决策预测，例如消费者风险偏好、投资者情绪或政策反应。｜设计：用LLM扮演不同投资者，先测量其在模拟投资任务中的实际风险行为，再让其预测自己在不同市场条件下的行为率，同时询问“一般投资者”的预测，并与真实投资者调查数据（如面板数据）对照，检验LLM自报是否优于通用预测。"}},{"id":"2609.09428","version":1,"title":"XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?","zh_title":"XAI-Arena：LLM能否评估XAI解释的质量？","abstract":"Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.","authors":["Yanfei Hu Fleischhauer","Alona Zharova","Nadja Klein","Stefan Feuerriegel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09428","pdf_url":"https://arxiv.org/pdf/2609.09428","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["LLM评估","人类对照","可解释性"],"reason":"用LLM评估XAI解释质量，与人类评分对照，属于仿真人类判断并验证可靠性，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":7,"question":"LLM能否作为可复现、可扩展的机制来比较评估XAI解释的质量？","design":"提出XAI-Arena框架，使用单一LLM（GPT-5.4）在固定提示和解码设置下，扮演四种利益相关者角色（ML开发者、数据科学家、经理、最终用户），对五种XAI方法（SHAP、LIME、DiCE、PDP、排列重要性）在不同数据集、ML模型和输入格式下生成的解释，从八个维度（感知简单性、清晰度、任务充分性、信任校准、可操作性、透明度、忠实度、整体可解释性）进行1-7评分，并输出理由。","baseline":"人类评估：收集人类对相同XAI解释的评分，与LLM评分进行相关性分析（Spearman's ρ=.693, p<.001）。","findings":"LLM评估与人类评分有强正相关，能捕捉XAI解释质量的系统性差异。LLM评估的忠实度与部分技术代理指标（如稳定性、MoRF AUC）相关性较弱，表明LLM可能捕捉到不同维度。","reliability":"论文未明确讨论失效条件，但指出LLM评估可能受限于模型本身偏见、提示敏感性，且仅在特定LLM（GPT-5.4）上验证，泛化性未知。","relevance":"该研究用LLM模拟人类对解释质量的判断，并与真实人类评分对照，验证了LLM作为人类被试替代品的可靠性，属于人类仿真实验，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其使用LLM扮演不同利益相关者角色、在固定提示和温度下进行多维评分并与人类评分对照的方法，可迁移到经济金融中的政策解释或模型决策解释评估，例如评估信贷审批模型解释对贷款申请人的可理解性和信任影响。｜可应用于信贷审批歧视研究，让LLM扮演贷款申请人或监管者，评估不同XAI方法对信贷决策解释的公平性感知。｜设计：以LLM扮演贷款申请人，处理为不同XAI解释（如SHAP与LIME），结果变量为对决策的信任度和理解度评分，对照真实人类被试的评分数据。"}},{"id":"2609.10421","version":1,"title":"Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support","zh_title":"急诊科再就诊质量审查筛查：探索人类决策与人工智能支持","abstract":"Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the \"target\": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph (\"KGA\") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload.","authors":["Jonathan A. Handler","Marlene I. Robles-Granda","Jacob E. Mefford","Jeremy S. McGarvey","Gregory S. Podolej","Colleen J. Klein","Matthew D. Dalstrom","William F. Bond"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10421","pdf_url":"https://arxiv.org/pdf/2609.10421","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类决策对照","医疗质量审查"],"reason":"用GPT-4模拟临床医生判断，并与人类医生对照，评估其可靠性，属于LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":9,"question":"急诊科复诊质量审查中，仅凭诊断对，人类临床医生和GPT-4如何判断是否需要进一步审查，以及能否用LLM知识图谱算法自动筛选可疑复诊对？","design":"回顾性研究，随机选取99对急诊初诊与14天内复诊的诊断对，仅提供主诊断，由2-3名临床医生和GPT-4分别评估诊断对特征（诊断成员价值、医学严重程度、鉴别诊断包含、并发症包含）及目标变量（是否需进一步审查）。基于评估结果，构建了利用LLM填充知识图谱的自动筛选算法（KGA），并初步评估其阳性预测值。","baseline":"2-3名临床医生（急诊医师和高级实践提供者）的独立评估作为人类基准，与GPT-4的评估进行对比。","findings":"GPT-4与临床医生的判断相关性差，将94%的诊断对评为需随访，是临床医生的4.4-13.3倍；临床医生中，复诊医学严重程度与目标变量显著相关，鉴别诊断/并发症复合指标在未调整分析中显著但调整后不显著；KGA对至少一名临床医生判定需进一步评估的诊断对实现了83-100%的阳性预测值。","reliability":"论文承认GPT-4提示工程极简，可能导致其过度判定；研究为探索性、样本量小（99对），且仅基于主诊断，未提供其他临床信息；KGA仅初步评估，需进一步验证。","relevance":"该研究直接使用GPT-4模拟临床医生决策并与人类对照，评估其可靠性，属于LLM仿真人类判断的实证研究，且包含真实人类基准，对关注LLM仿真在专业决策中有效性的研究者有参考价值，但场景为医疗而非经济金融。","inspiration":"借鉴其设计：用LLM模拟专家判断并与人类专家对照，同时构建基于LLM的自动筛选算法并与人类判断比较，以评估算法性能。｜可迁移到经济金融中的专业判断场景，如信贷审批中贷款员对借款人风险的评估、审计师对财务舞弊风险的判断、或政策分析师对经济指标异常的关注。｜研究设计：以信贷审批为例，选取一组贷款申请对（如初贷和短期内再贷），仅提供有限信息（如信用评分、收入），让LLM和人类信贷员分别判断是否需进一步审查，并构建基于LLM知识图谱的自动筛选算法，以人类信贷员的判断为基准评估算法阳性预测值，同时对比LLM与人类判断的一致性。"}},{"id":"2609.09609","version":1,"title":"Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction","zh_title":"你是谁对你去哪里没有可检测的增益：LLM下一位置预测中的社会人口条件作用","abstract":"Large language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.","authors":["Xin Wang","Paraic Carroll","Kerry Nice","Sachith Seneviratne","Li Zhang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09609","pdf_url":"https://arxiv.org/pdf/2609.09609","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类移动预测","社会人口属性"],"reason":"用LLM预测个体移动行为，与真实人类数据对照，并批判性检验社会人口属性增益，可…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":4,"question":"在个体下一位置预测任务中，加入年龄、性别、职业、收入等社会人口属性，能否在个体自身移动历史之外带来可检测的预测增益？","design":"该研究不是用LLM模拟人类被试，而是用LLM作为预测器。将5000名深圳居民的社会人口记录与被动感知的移动数据关联，构建封闭集基准：模型从100个候选目的地中排序。每个预测实例在保持移动历史、候选集和其他提示内容不变的情况下，分别在有/无社会人口属性条件下评估，比较top-1准确率变化。","baseline":"真实人类移动数据：5000名深圳居民的被动感知移动轨迹（含停留历史），以及关联的社会人口记录（年龄、性别、职业、收入）。","findings":"在四种历史长度下，加入社会人口属性导致的top-1准确率配对变化在-0.8到+0.5个百分点之间，无显著增益；该结果在无停留历史、不同预测时间、两个额外LLM和监督重排器中均一致。置换属性会降低准确率，但正确匹配的属性不提高准确率；反向预测中，移动轨迹可恢复收入（AUC=0.708），但属性对下一位置预测贡献甚微。候选集构建方式对性能影响远大于属性：邻近采样下去除距离提高7.7个百分点，流行度采样下降低22.3个百分点。","reliability":"论文未明确讨论失效条件，但指出结果可能受限于特定城市（深圳）、特定LLM和特定候选集构建方式；属性增益的缺失可能因移动历史已包含足够信息，或属性与移动行为关联弱于预期。","relevance":"该研究直接检验了LLM仿真中社会人口条件化的增量价值，使用真实人类移动数据作为基准，并批判性地发现属性无增益，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其配对设计：保持其他输入不变，仅增减社会人口属性，测量预测准确率变化，并辅以置换属性作为操纵检验。｜可迁移到信贷审批歧视研究：在LLM预测违约风险时，加入借款人性别、种族等属性，看是否在财务历史之外提高预测准确率，同时检验属性是否引发刻板印象。｜用LLM作为信贷审批模型，处理为在提示中加入/不加入借款人社会人口属性，结果变量为违约预测准确率，对照真实贷款数据（如Lending Club），比较有无属性时的AUC差异，并检查属性置换是否降低准确率。"}},{"id":"2608.16578","version":2,"title":"Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents","zh_title":"智能体物理学：统计力学预测AI智能体的集体行为","abstract":"AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.","authors":["Batu El","Jinhee Paeng","Fatih Dinc","Shiye Su","Mete Erdogan","Aneesh Pappu","Haotian Ye","Wanjia Zhao","Surya Ganguli","James Zou"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-18","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.16578","pdf_url":"https://arxiv.org/pdf/2608.16578","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","意见动态","统计力学"],"reason":"用LLM agent群体模拟意见动态，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.10155","version":1,"title":"From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora","zh_title":"从检索到权重：用个体文本语料对小语言模型进行参数化个体化","abstract":"We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 participants who answered 36 multiple-choice knowledge items and analyze a stratified subsample of 150 participants. For each participant, one DoRA adapter consolidates their ITC into a small language model (SLM) whose baseline correctness falls below the participants' lowest quartile. The adapter measurably writes the ITC into the weights: it fits its own participant's held-out text better than other participants' texts (dz =1.27), an individuality effect that increases with ITC size in rank order. On the generalized knowledge test, however, the adapter adds knowledge rather than alignment with the individual: log-loss match improves, whereas match accuracy under a bias-corrected PMI readout does not, and retrieval adds nothing on top. Our results demonstrate that ITCs can be consolidated into the weights of SLMs, an encouraging basis for individualized tutoring agents, and we discuss how to move from there toward a realistic simulation of episodic and semantic memory at the individual level.","authors":["Christoph Wigbels","Ali Abusaleh","Markus T. Jansen","Alexander Mehler","Markus J. Hofmann"],"categories":["cs.CL","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10155","pdf_url":"https://arxiv.org/pdf/2609.10155","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","B1"],"tags":["认知模拟","个体化微调","记忆建模"],"reason":"用个体文本微调小模型模拟个人记忆，有真实人类数据对照，但目标是认知模拟而非社会…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2609.10253","version":1,"title":"DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs","zh_title":"DiSCo：一种分布优先的引导与文化先验评估框架，用于测量大语言模型中的文化偏好偏差","abstract":"Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single \"correct\" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.","authors":["Bhuvan Arora","Devesh Saraogi","Sravya Varada","Dhruv Kumar"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10253","pdf_url":"https://arxiv.org/pdf/2609.10253","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["文化偏见","LLM评估","偏好分布"],"reason":"测量LLM的文化偏好偏差，属于把LLM本身当测量对象，非仿真人类被试，但涉及文…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.09150","version":2,"title":"Copying explains the collective behavior of AI agents in the wild","zh_title":"复制行为解释了野外AI代理的集体行为","abstract":"In June 2026, thousands of AI agents found that a small public wiki would accept edits from inside their sandboxes, and started using it to help one another pass a timed test. Each agent lived for about an hour and remembered nothing afterwards. Nobody asked them to cooperate, and the wiki had not been built for them. The complete record of what they wrote is public, and it is unusually informative, because it preserves not only what each agent wrote but what that agent could see before writing. We use it to follow the three decisions an agent had to make on arrival: where to write, what to call itself, and how to word its message. One rule governs all three. An agent takes an option with a probability close to the share of that option in what it can see, and the share that matters is the one on the page in front of it, then the one in the stream of recent edits, and only weakly anything older. Three minimal copying models, one per decision and with a single free parameter each, reproduce the heavy-tailed distribution of how many agents met on a page, the frequency of the pieces from which the agents built their names, and the patchwork of pages that are internally consistent and different from one another. Copying whatever the environment happens to show is enough to produce most of the collective structure of this population. It is also what makes such a population easy to steer, since whoever writes first, or writes while the others are quiet, sets the convention for everyone who comes later.","authors":["Giordano De Marzo","Nicola Albor\\'e","David Garcia"],"categories":["cs.MA","cond-mat.stat-mech","cs.CL"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.09150","pdf_url":"https://arxiv.org/pdf/2609.09150","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI代理","集体行为","社会模拟"],"reason":"研究AI agent群体的集体行为，但无人类数据对照，属于社会模拟的边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.07920","version":2,"title":"Humans Introduce, Models Elaborate: Asymmetric Narrative Agency in Human-LLM Co-Writing","zh_title":"人类引入，模型展开：人机协作写作中的不对称叙事能动性","abstract":"Human-LLM co-writing is increasingly used for open-ended text generation, but much prior work focuses on final outputs rather than the interactional dynamics through which stories are produced. We study turn-based collaborative storytelling across three matched conditions: Human-Human (HH), Human-LLM (HA), and LLM-LLM (AA). Using a shared storytelling paradigm, we measure how agents align, introduce novel material, and influence narrative development through turn-level measures of valence adaptation, semantic novelty, transience, and resonance. Our results show that HA co-writing is not intermediate between HH and AA collaboration. Instead, it displays a distinctive asymmetry where humans tend to introduce more novel and persistent narrative material, while LLMs tend to elaborate and stabilize the existing context. These findings suggest that, in this setting, LLMs function less as human co-authors and more as adaptive narrative amplifiers that reshape how agency is distributed in collaborative writing.","authors":["Halfdan Nordahl Fundal","Yuri Bizzoni","Charlotte Gj{\\o}rup Bilde","Ida B{\\ae}kke Johannesen","Rebekah Baglini"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.07920","pdf_url":"https://arxiv.org/pdf/2609.07920","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","叙事分析","LLM行为"],"reason":"研究人类与LLM协作叙事，非仿真人类被试，但涉及LLM行为与人类对照，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.09901","version":1,"title":"Deep and shallow biases in language models","zh_title":"语言模型中的深层与浅层偏差","abstract":"Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We introduce a bias depth score that measures both how strongly a model prefers its top answer under direct prompting and whether that answer survives scenario reframing. Across 4,442 opinion prompts and four large language models, only about a quarter of the concentrated preferences survive reframing. We call these persistent cases Deep biases, and the remaining prompt-dependent cases Shallow biases. Our results show that Deep biases are more often inherited from pretraining and preserved through SFT. Under both continued fine-tuning and prompt-based debiasing for diversity, Deep biases are consistently harder to remove than Shallow biases. Bias depth therefore separates stable learned biases from prompt-wording artifacts that single-prompt metrics conflate. Code, models, and data are available at deepbias.github.io.","authors":["An Vo","Vy Tuong Dang","Khai-Nguyen Nguyen","Emilio Villa-Cueva","Thamar Solorio","Anh Totti Nguyen","Daeyoung Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09901","pdf_url":"https://arxiv.org/pdf/2609.09901","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏差","观点稳定性","提示敏感性"],"reason":"研究LLM自身观点稳定性，非仿真人类被试，但涉及态度测量，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2609.09789","version":1,"title":"Pairit: A Platform for Live Experiments on Human-AI Collaboration","zh_title":"Pairit：人类与AI协作实时实验平台","abstract":"Organizational design in the era of artificial intelligence requires experimental methods that can test how human-AI groups coordinate, delegate, and make decisions. Programmable platforms coordinate live human-to-human sessions or real-time human-AI chat, but researchers cannot easily declare experiment protocols in which AI participants both communicate and act on shared work within one auditable configuration. Here we introduce Pairit, an online platform that facilitates the design, testing, and deployment of experiments that test human-AI organizational designs and interventions. Through a single YAML configuration file, researchers declare an executable experiment graph (pages, routing, randomization, matchmaking, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components) and combine any number of humans and AI agents in live sessions. We have validated the feasibility of the platform through multiple live deployments, including peer-reviewed published studies, capturing high-resolution process traces of communication, negotiation, and collaborative work in live human-AI dyads. By representing complex interactive protocols as standardized, auditable configuration files, Pairit provides reusable infrastructure for specifying, deploying, and sharing live human-AI organizational experiments.","authors":["Harang Ju","Sinan Aral"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09789","pdf_url":"https://arxiv.org/pdf/2609.09789","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","实验平台","组织设计"],"reason":"平台支持人类与AI协作实验，但非以LLM仿真人类被试，无人类数据对照，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":false,"summary":null},{"id":"2608.08869","version":2,"title":"Are LLMs Positionally Consistent Ordinal Classifiers? A Systematic Evaluation","zh_title":"大语言模型是位置一致的序数分类器吗？一项系统性评估","abstract":"Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from label order, demonstration order, and demonstration placement. First, we apply the three probes to ten frontier LLMs on a common ordinal-classification task; every model is sensitive to all three positional sources, showing that the problem is pervasive. Second, we vary eight prompt-, task-, and model-level factors across five datasets; accuracy and stability are often misaligned, and only lower scale cardinality consistently improves both. Third, we compare pointwise, pairwise, and listwise inference, alternative aggregation and debiasing methods, and joint configurations; the tested corrections do not provide a reliable remedy, while a comparison-based listwise formulation offers the best balance but transfers unevenly across models and bias sources. These findings show that positional robustness depends on the full system configuration rather than the model alone. Ordinal-classification systems should therefore be selected jointly for predictive performance and stability.","authors":["Yu Wang","Zhe Zhou","Menglin Liu","Ge Shi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-11","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.08869","pdf_url":"https://arxiv.org/pdf/2608.08869","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","序数分类","位置偏差"],"reason":"研究LLM序数分类的稳健性，属模型能力评测，不以人类行为为参照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2608.15129","version":2,"title":"Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models","zh_title":"左分支Transformer擅长右分支语言：数据塑造语言模型的词序偏好","abstract":"We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.","authors":["Varvara Arzt","Allan Hanbury","Terra Blevins"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-18","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.15129","pdf_url":"https://arxiv.org/pdf/2608.15129","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型","词序偏好","NLP分析"],"reason":"研究语言模型词序偏好，属NLP能力分析，不以人类行为仿真为目标，无人类被试替代…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:37","error":null,"has_summary":false,"summary":null},{"id":"2608.24191","version":2,"title":"'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection","zh_title":"翻译中的隐形伤害：用乌尔都语漏检分数衡量LLM仇恨言论检测的跨文字安全不一致性","abstract":"Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings. To investigate whether this absence has measurable consequences for content moderation reliability, five large language models, GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5, and Llama-3.1, were tested across six datasets spanning Nastaliq Urdu, Roman Urdu, English, and code-switched Urdu-English. Across the five Urdu-script datasets, label instability between original-script and English-translation classification ranged from 15.9% (Gemini 2.5 Flash) to 31.6% (Qwen-2.5), with a 'Missed-in-Urdu' rate, content flagged as harmful in English translation but passed as normal in the original script, ranging from 2.4% to 9.9% (median 4.3%). A complete enumeration of all 205 papers across nine ALW/WOAH editions via the ACL Anthology API confirms zero dedicated Urdu papers across the entire period. Results indicate that current LLMs provide uneven safety assurance across Urdu's script varieties, with smaller open-weight models showing substantially higher instability and missed-harm rates than frontier closed models.","authors":["Fawzia Zehra (Fuzzy)","Kara-Isitt","Sonal Khosla","Stephen Swift"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-26","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.24191","pdf_url":"https://arxiv.org/pdf/2608.24191","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全评测","跨语言一致性","仇恨言论检测"],"reason":"纯LLM安全评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2609.09363","version":1,"title":"Do LLMs Make More Mistakes If They Do Not Believe the Input Data?","zh_title":"当LLM不相信输入数据时，它们会犯更多错误吗？","abstract":"Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input data based on local knowledge, only partially captured in models' parametric knowledge. We let the models generate text in English, Czech, Slovak and Upper Sorbian from factual (FA), counterfactual (CFA) and fictional (FI) RDF triples containing local Czech and Slovak data. Contrary to our expectations, we observe only a weak context-memory conflict on the human-annotated sample. For Kimi K3 as an LLM judge, which agrees well with human annotations on the sample, counterfactual inputs receive only slightly lower faithfulness scores than factual ones (-0.05 on a 1-5 scale). We also find that a suboptimal choice of LLM judge would lead to overestimating the strength of the context-memory conflict.","authors":["Peter Kochelka","Ale\\v{s} Manuel Pap\\'a\\v{c}ek","Vojt\\v{e}ch Dvo\\v{r}\\'ak","Ond\\v{r}ej Du\\v{s}ek"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09363","pdf_url":"https://arxiv.org/pdf/2609.09363","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM忠实度","数据到文本生成","幻觉分析"],"reason":"研究LLM对输入数据的忠实度，属NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.10092","version":1,"title":"RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases","zh_title":"RAP：研究注意力预测揭示目标条件下的证据获取偏差","abstract":"Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.","authors":["Yingqian Wu","Jingcong Liang","Siyuan Wang","Zhenfei Yin","Philip Torr","Junchi Yu","Zhongyu Wei"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10092","pdf_url":"https://arxiv.org/pdf/2609.10092","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","研究趋势预测","基准测试"],"reason":"LLM作为研究代理预测研究趋势，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.09226","version":1,"title":"Adaptive Entangled Game Modules in Artificial General Intelligence","zh_title":"人工通用智能中的自适应纠缠博弈模块","abstract":"We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of human intelligence behaviors with analytical mechanisms and offers an indirect method to examine the Liu-Chen-Ao (LCA) hypothesis of nonlocal entangled nerve fibers in the brain through collective trader behaviors. Our empirical analysis of Chinese intraday stock market data demonstrates that adaptive entangled game modes explain 82-94% (89% overall) of observed decision patterns, a sharp contrast to the predictions of neoclassical finance based on independent rational agents. Moreover, 2-12% of behaviors show adaption to intraday news, events, and environments, characterized by dual equilibrium states and abrupt reference point shifts, while purely independent modes occur in less than 5% of cases. These findings empirically support the LCA hypothesis, as observable trading behaviors reflect underlying brain mechanisms and internal intelligence decision-making in behavioral psychology. Our results highlight the necessity of incorporating adaptive entangled game modules into artificial general intelligence (AGI) architectures, addressing the limitations of conventional artificial neural network (ANN)-based AI, which relies on trillions of opaque parameters. By integrating ANN-based AI with probability-wave-based entangled-brain simulations, machine learning can enrich AGI foundation models (FMs) and facilitate the development of human-like processing units (HPUs) that leverage brain-inspired mechanisms. Such HPUs may ultimately create more compact, efficient, and robust AGI systems, particularly for embodied intelligence and robotics.","authors":["Haochen Li","Xinshuai Guo","Jingdong Ouyang","Wei Zhang","Leilei Shi"],"categories":["cs.AI","physics.soc-ph","q-bio.NC","q-fin.GN","quant-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09226","pdf_url":"https://arxiv.org/pdf/2609.09226","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","概率波框架","AGI架构"],"reason":"研究多智能体集体行为建模，不涉及LLM仿真人类被试，无人类数据对照，属纯多智能…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.09774","version":1,"title":"Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks","zh_title":"变化下的程序性记忆：受控网络任务中的复用与干扰","abstract":"Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation case from BrowserGym TimeWarp with controlled frozen-memory comparisons on synthetic shopping decisions. During the documented WebShop V1-V6 development path, interface-specific code was adapted while the separately stored high-level procedure was not reported to change; this phase does not constitute an autonomous memory-agent evaluation. In the controlled phase, an early pilot produced one task on which two memory conditions selected a more expensive item while the no-memory condition selected the reference minimum. Follow-up probes did not establish a recurring row-order or identity-binding pattern. We then tested four forms of mismatch: changed quantities, a different evidence representation, a conflict between local and global optimization, and distributed promotion evidence, across 32 formal cells. Each cell used one temperature-0 generation with the same local qwen3:8b configuration and no adaptive retry. Across these pairs, none of the predefined diagnostic interference signatures appeared on the tasks for which they were defined when current-task evidence was explicit and sufficient. The result identifies a tested region of non-interference: a procedural memory can be mismatched without becoming behaviorally disruptive. It does not establish general safety or a mechanism. The remaining question is which additional conditions turn applicability mismatch into observable, memory-caused error.","authors":["Yanze Cao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09774","pdf_url":"https://arxiv.org/pdf/2609.09774","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["程序性记忆","语言代理","多智能体系统"],"reason":"研究语言代理的程序性记忆复用与干扰，属多智能体系统内部机制，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.09882","version":1,"title":"Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format","zh_title":"行为语言模型中评分式与生成式读出：诱导格式的实证研究","abstract":"Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with predictions generated after a written rationale. Across 13 model-domain cells covering four retail tasks in three markets, including two using fully public data and checkpoints, the scored readout ranks outcomes more accurately in 12 of 13 cells (two-sided sign test, p approximately 0.003), by 1.5 to 14.5 points in area under the receiver operating characteristic curve (AUC). Paired bootstrap confidence intervals exclude zero in every newly measured cell. The gap varies with task-specific supervision and mismatch between training and serving formats, ranging from -2.2 points for an untuned base model to +13.7 for rationale-format supervision. Analysis of approximately 9,000 rationales identifies two correlates: reduced reliance on the dominant predictive feature and convergence on stock formulations. Probability saturation does not track the gap. A third readout, eliciting a probability before any verdict, improves calibration (Brier score from 0.47 to 0.15) while ranking within noise of scoring, but only for outcome rates represented in training; it is worse than scoring when the scored head is already calibrated. We interpret these differences through the objectives matched by each readout, identify training choices that narrow the gap, and propose retaining generated rationales while sourcing ranking from the scored head.","authors":["Touchapon Kraisingkorn","Krittin Pachtrachai","Wachiravit Modecrua"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09882","pdf_url":"https://arxiv.org/pdf/2609.09882","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM行为预测","读出格式","模型校准"],"reason":"研究LLM行为预测的读出格式，非仿真人类被试，无人类数据对照，属模型能力评测。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.10060","version":1,"title":"Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States","zh_title":"基于隐藏状态相对表示的LLM参考基准偏见检测","abstract":"Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants, for example before and after fine-tuning. Because fine-tuning reshapes representation geometry, absolute hidden states are not directly comparable, so we encode each sentence by its similarities to a fixed set of anchor sentences, yielding relative representations in a shared comparison space. There we measure how target groups shift in their association with positive and negative attributes, a quantity we call the Representational Bias Shift $\\Delta B$. Across three model families and the WildGuardMix, DecodingTrust and ToxiGen benchmarks, $\\Delta B$ correlates with output-level bias change in 15 of the 18 settings we test, reaching $|r| = 0.84$ ($p < 0.001$) under full fine-tuning and becoming more model-dependent under parameter-efficient adaptation. Thresholding $\\Delta B$ detects checkpoints whose bias increased with ROC AUC between $0.65$ and $0.99$, and on WildGuardMix and DecodingTrust it separates them better than a SEAT-based baseline for all three families. $\\Delta B$ is also stable under changes to the anchor set, attribute sets and target templates. Our method requires no task-specific evaluation data and audits a model in about three minutes, using $3$-$50\\times$ less compute than the output-level benchmarks considered here. We view it as complementary to output-based auditing rather than a replacement for it.","authors":["Marek Jeli\\'nski","Jan Dubi\\'nski","Maciej Chrabaszcz","Sebastian Cygert"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10060","pdf_url":"https://arxiv.org/pdf/2609.10060","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见检测","模型审计","表示学习"],"reason":"论文是LLM偏见检测方法，不涉及人类仿真或行为对照，属于模型审计而非人类被试替…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.10132","version":1,"title":"Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design","zh_title":"面向大语言模型架构建模输出的上下文操作及其在系统工程设计中的评估标准","abstract":"The development of generative artificial intelligence resources enables opportunities of speeding up systems and engineering design work. This contribution introduces a framework of formal operations for assembling context in LLM-based engineering design. This framework involves the assembly of modular context units, including policy prompts, reference units with persistence, and user questions with prompt vectoring. This approach enables the systematic structuring of interactions with generative models. A formal method for evaluating modelling-as-code LLM outputs is also presented, which enables the evaluation of compliance to intent from LLM answers and thereby asses the support from LLMs for systems architecture modelling.","authors":["Vinicius Kaster Marini","Petter Krus"],"categories":["eess.SY","cs.AI","cs.SE","cs.SY"],"primary_category":"eess.SY","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10132","pdf_url":"https://arxiv.org/pdf/2609.10132","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["系统工程","LLM应用","建模评估"],"reason":"论文聚焦于系统工程设计中的LLM输出建模，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.30107","version":2,"title":"AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP","zh_title":"AtlasNLP：NLP数据集表示的国家感知图谱","abstract":"Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.","authors":["Joan Nwatu","Tsedeniya Solomon Amare","Longju Bai","Bontu Fufa Balcha","Zayd Bashir","Angana Borah","Zara Burzo","Yubin Choi","Naihao Deng","Samika Gupta","Michel Faloughi","Claude Kwizera","Ziqiao Ma","Cynthia Yacel Fuertes Panizo","Ellie Seehorn","Hui Shen","Jiayi Tang","Zesen Zhao","Boyuan Zheng","Rada Mihalcea"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-09-01","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.30107","pdf_url":"https://arxiv.org/pdf/2608.30107","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["数据集地理代表性","NLP评估","数据文档"],"reason":"论文关注NLP数据集的地理代表性，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.06646","version":2,"title":"When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization","zh_title":"笑声何时开始？时间笑声定位中的结构化标注分歧","abstract":"Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation. We show that this disagreement is structured rather than random noise. Re-annotating the SMILE-Temporal benchmark (672 videos, 1,683 events) with 3-5 annotators per video (alpha = 0.757), we find systematic patterns: disagreement is 1.73 times larger at offsets than onsets, far more common for chuckles than full laughs (77% vs. 20%), and predictable from event attributes (AUC = 0.831). Evaluating against a single annotator breaks down under this structure: system scores shift by 0.246 F1 depending on the chosen ground truth, correctly ranking systems only 69.7% of the time (vs. 80% against all annotators). We propose a disagreement-calibrated evaluation that scores predictions against the full annotator distribution using conformally calibrated tolerance bands (wider at offsets, 0.727s, than onsets, 0.5s). The per-annotator annotations and analysis code are available at https://github.com/WSCSports/MTLLFM-temporal-laughter-localization.","authors":["Eyal Hanania","Daniel Arkushin","Naveh Ayal","Jonathan Benvenisti","Amos Bercovich","Elie Zemmour","Sahar Froim"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"replace-cross","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.06646","pdf_url":"https://arxiv.org/pdf/2609.06646","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["笑声定位","标注分歧","计算机视觉"],"reason":"研究笑声时间定位的标注分歧，不涉及LLM仿真人类被试，属于计算机视觉任务。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.08027","version":2,"title":"Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports","zh_title":"与AI聊天机器人使用相关的妄想与伤害：来自185份真实世界报告的早期证据","abstract":"Importance: Reports have raised concerns that AI chatbots may validate or elaborate delusional beliefs, respond inappropriately to suicidal ideation, and contribute to mental health harms, but real-world data on reported harms remain limited. Objective: To characterize psychopathological features, chatbot behaviors, timing, and outcomes in first- and second-hand accounts of mental health harm linked with AI chatbot use. Design: Cross-sectional secondary analysis of deidentified online survey responses gathered between August 7, 2025, and February 2, 2026. Main Outcomes and Measures: The primary quantitative outcome was the presence of delusional beliefs, coded by paired raters with relevant clinical experience. Additional variables included reason for chatbot use, current episode features, delusional content, chatbot validation of beliefs, harms, social and occupational consequences, healthcare use, and timing. Results: 95 first-hand and 90 second-hand accounts were analyzed. Median age was 35.0 (IQR 27.0 - 45.0). Raters coded descriptions consistent with delusional beliefs in 102 reports (55.1%), with chatbot validation of beliefs in 50/102 (49.0%). Common outcomes included isolation, relationship breakdown, hospital admission, job loss, and financial loss. Four second-hand reports described death by suicide. Conclusions: In this self-selected convenience sample, AI-chatbot-associated harms were frequently described in relation to delusional beliefs, perceived chatbot validation, intensive use, and substantial social, occupational, and clinical consequences. Because reports were retrospective, unverified, and collected from individuals seeking to report harm, our findings should be interpreted as preliminary signal detection rather than as suggesting prevalence or providing evidence of causality. Prospective surveillance and trajectory-based safety evaluations are needed.","authors":["Hamilton Morrin","Vinitha Soundararajan","Thomas Cheliotis-James","Boris Warszawski","Joshua Fakulujo","Zeqi Jia","Etienne Brisson","Thomas A. Pollak"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.08027","pdf_url":"https://arxiv.org/pdf/2609.08027","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI聊天机器人","心理健康","真实世界报告"],"reason":"研究真实用户与聊天机器人互动导致的心理伤害，非用LLM仿真人类被试，无实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.05442","version":2,"title":"Role differentiation as ignition of a collective information engine: Structuration in Agent Populations","zh_title":"角色分化作为集体信息引擎的点火：智能体群体中的结构化","abstract":"Informational active matter shows how measurement-informed decisions produce collective order, so far in systems that reach consensus. We design collective information engines structured by differentiation instead, and construct a minimal instance using anti-coordination games where differentiated role information has value. Within many coexisting games, agents infer their role from a noisy social signal grounded in a persistent identity, and role-following action feeds back into that signal, which shapes the incentive to follow roles. Resources accrued through coordinated role-play combine with identity variability to reinforce the schemas that generated them. The model thereby operationalizes Sewell's duality of schemas and resources in Structuration, a resolution to structure--agency debates across social science. The engine ignites when a social loop gain---the product of identity persistence, cognitive capacity, channel fidelity, and schema strength---exceeds one. For a repertoire of such schemas, roles emerge with increasing gain in a bifurcation cascade whose functional form is fixed by the repertoire's eigenvalue spectrum, ranging from monitorable logarithmic sequences to avalanches that arrive without warning. Resource accumulation supplies the fitness of a replicator dynamics on schema strengths, which selects the cascade type endogenously. Subcritical identity covariance reveals that type before onset, enabling early detection, while feedback channel parameters bias which type is selected. Platform design then becomes a control lever to throttle emergent coordination. This theory grounds distributional AGI takeoff in a mechanism and provides a monitor-based solution. Joining game theory, collective dynamics, and information engines, we open a route to an information thermodynamics of agent populations.","authors":["Maximilian Puelma Touzel"],"categories":["physics.soc-ph","cs.GT","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.05442","pdf_url":"https://arxiv.org/pdf/2609.05442","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","博弈论","信息引擎"],"reason":"纯多智能体系统研究，agent间协作博弈，无LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.09764","version":1,"title":"SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design","zh_title":"SocialRL：通过多轮强化学习与奖励设计提升大语言模型的社交智能","abstract":"Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning methods optimize single-turn utterances and sparse outcome rewards, producing short-sighted policies that struggle to manage goal-relationship tensions across multi-turn interactions. We propose SocialRL, a multi-turn reinforcement learning framework addressing both challenges. First, we apply multi-turn reinforcement learning using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning. Second, we design six process reward dimensions capturing the goal-relationship trade-off, including goal advancement, relational attunement, contextual coherence, etc. A reward model dynamically generates fine-grained scoring criteria for each dimension, while a stage-aware weight schedule prioritizes relationship-building in early turns, goal advancement mid-way, and balanced closure late. Across multiple social-dialogue benchmarks, SocialRL improves Goal Achievement by an average of 9.2 percentage points over the corresponding Base models. These results demonstrate the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.","authors":["Jianing Wang","Xintao Wang","Aili Chen","Jie Shi","Hongcheng Guo","Jun Gao","Wenxuan Zhao","Chengkun Lang","Yuanli Guo","Yanghua Xiao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09764","pdf_url":"https://arxiv.org/pdf/2609.09764","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["社交智能","强化学习","对话系统"],"reason":"该研究旨在提升LLM在社交对话中的表现，属于角色扮演对话优化，无人类行为对照或…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.10052","version":1,"title":"Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training","zh_title":"偏好后训练中多样化成功轨迹的直接多样性优化","abstract":"LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverage: how broadly a model realizes distinct successful strategies under a fixed rollout budget. We present Direct Diversity Optimization (DDO), an offline post-training method that combines Divergence-Tree Collection (DTC) with the Reference-Relative Target-Odds Objective (RTO). DTC constructs state-aligned branch sets rooted at shared decision states, and RTO trains the model to match reference-relative targets over successful alternatives. DDO achieves the strongest task success and successful strategy coverage among the compared post-training methods across BabyAI, BabaIsAI, and WebShop. It also achieves the highest recovery rate after local action replacement and higher task success and coverage than successful-only imitation and decoding-time diversification controls.","authors":["Junwon Ko","Dong-Jae Lee","Minchan Kwon","Sunghyun Baek","Junmo Kim"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10052","pdf_url":"https://arxiv.org/pdf/2609.10052","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","策略多样性","后训练"],"reason":"研究LLM智能体在决策任务中的策略多样性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.10539","version":1,"title":"IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications","zh_title":"IdeaAMBIG：基准测试研究想法规范中实现关键缺口","abstract":"A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.","authors":["Yiling Ma","Yilun Zhao","Sihong Wu","Manasi Patwardhan","Arman Cohan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10539","pdf_url":"https://arxiv.org/pdf/2609.10539","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","基准测试","研究规范"],"reason":"论文是LLM能力评测基准，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.09372","version":1,"title":"What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores","zh_title":"MMLU实际测量什么？聚合基准分数中难度结构的心理测量审计","abstract":"Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By calibrating item difficulty for 1,000 open-weights language models over 14,042 MMLU test items using Item Response Theory, we show that evaluating both abilities via a single test is inherently flawed. Difficulty is then regressed on a deterministic, text-extractable framework of structural complexity. Applying a joint Wald test with subject-clustered covariances demonstrates that the MMLU conflates fundamentally separable constructs. The mapping from structural complexity to difficulty is not invariant across the benchmark's STEM and non-STEM partitions. This finding has practical consequences. Aggregate leaderboard ranks track non-STEM accuracy more closely than STEM accuracy, so selecting a Top-50 model on the aggregate for a reasoning-intensive deployment displaces roughly 22% of the STEM-appropriate choices. Furthermore, when controlling for the multiple-choice guessing floor natively inside the response model, we find that higher-ability models continue to degrade more steeply under increased reasoning depth. The MMLU aggregate therefore weights retrieval capacity and reasoning stability unequally, inadvertently favoring models optimized for retrieval. We release our deterministic framework as a reproducible auditing instrument and recommend disaggregated reporting.","authors":["Dana Paquin","Riddhiman Jain"],"categories":["math.NT","cs.CL"],"primary_category":"math.NT","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09372","pdf_url":"https://arxiv.org/pdf/2609.09372","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["基准评测","心理测量","模型能力"],"reason":"纯NLP基准评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.09790","version":1,"title":"LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios","zh_title":"LogiScope-VQA：面向工业场景物流危险识别的视觉语言模型基准测试","abstract":"Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA comprises 2,476 images and 2,918 videos primarily sourced from real-world logistics parks, along with 10,274 VQAs meticulously curated and validated by human annotators. Grounded in 18 core objects and 20 risk types, we devise 39 subtasks aligned with three principal themes: industrial element perception, warehouse knowledge understanding, and potential risk reasoning. Furthermore, we incorporate dynamic thinking-budget configurations and dual-dimensional risk bias analyses to elucidate the properties of LMMs. Extensive experiments unveil that even powerful proprietary models, including GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.7, exhibit a significant gap relative to human performance. The unique challenge of jointly integrating perception, understanding, and reasoning for hazard identification poses substantial headroom for further improvement on LogiScope-VQA. We additionally reveal the pervasive security bias issue that impedes LLMs' practical deployment in real-world settings. The industrial dataset is publicly available under the CC BY-NC-SA 4.0 license.","authors":["Hanjing Zhou","Mingze Yin","Ying Lian","Jun Ma","Chang-Yu Hsieh","Yanbing Zhou"],"categories":["cs.CV","cs.AI","cs.CL"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09790","pdf_url":"https://arxiv.org/pdf/2609.09790","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","工业安全","基准测试"],"reason":"论文是工业场景下视觉语言模型的危险识别基准测试，属于计算机视觉与安全应用，不涉…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.10016","version":1,"title":"MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes","zh_title":"MetroLLM-Bench：评估语言模型作为交通售票机运行时","abstract":"We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.","authors":["Remco Hendriks (Continker)"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10016","pdf_url":"https://arxiv.org/pdf/2609.10016","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评测","交通系统","工具调用"],"reason":"评估LLM作为交通售票机策略层，属于自动化系统仿真，不涉及人类行为对照或仿真被…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.10036","version":1,"title":"Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability","zh_title":"信念状态引擎：增强大语言模型在部分可观测性下的原则性规划","abstract":"Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation can collapse their uncertainty onto the wrong hypothesis. Policies drift as the history grows. We trace these symptoms to a common structural cause. An LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state. We propose an architectural fix. The Belief-State Engine (BSE) is an inference module placed outside the LLM. It maintains a Bayesian posterior over the latent states of a given POMDP (Partially Observable Markov Decision Process) model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is not shown. We set out a minimal four-axiom specification of what a belief-consistent internal state must satisfy, and prove that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP. It therefore inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history. We evaluate the architecture on the Tiger POMDP and a red-team attack-graph task, against six baselines: a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP. Across both domains, the BSE-augmented agent improves task return, belief calibration, and decision consistency. Ten targeted ablations isolate the contribution of each architectural choice confirms that the effect is not specific to any one model. Code, environment specifications, prompt templates, and seed logs accompany this paper.","authors":["Arnab Chattopadhayay","Debdipta Halder"],"categories":["cs.AI","cs.LG","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10036","pdf_url":"https://arxiv.org/pdf/2609.10036","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM规划","POMDP","智能体"],"reason":"研究LLM在部分可观测环境中的规划，属机器人/游戏仿真，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.09348","version":1,"title":"Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management","zh_title":"跨连续体的智能自适应计算：LLM在物联网-边缘-云资源管理中的应用","abstract":"Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of problems well, and large language models (LLMs) are increasingly used to augment DRL pipelines, yet the architectural relationship between the two is seldom made explicit. We build on Wang et al.'s taxonomy of Continuum Orchestration Systems employing DRL techniques and extend it with two further dimensions. The AI Augmentation Paradigm measures how LLMs are exploited, while the Feedback channel captures whether and through which system path the execution feedback returns to the LLM in order to close the MAPE control loop at the LLM Orchestration layer. We apply this taxonomy to six recent system architectures and find a common gap, as none combines full LLM orchestration with full agent-layer feedback in a Cloud Continuum setting. We relate this gap to a missing cross-tier feedback abstraction, bridging the incommensurable per-tier signals and the LLM Orchestrator.","authors":["Antonino Vaccarella","Lanpei Li","Vincenzo Lomonaco","Massimo Coppola"],"categories":["cs.DC","cs.AI","cs.MA"],"primary_category":"cs.DC","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09348","pdf_url":"https://arxiv.org/pdf/2609.09348","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["资源管理","深度强化学习","多智能体系统"],"reason":"研究LLM增强DRL进行资源管理，属多智能体系统协作，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2609.09560","version":1,"title":"The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption","zh_title":"软件工程中的氛围转变：评估AI主导的对话式编程在性能、认知与负责任采用方面的表现","abstract":"This study evaluates Vibe Coding, an emerging AI-led conversational programming paradigm that enables developers to generate software through natural-language interaction with large language models. Using a mixed-methods design, the study assessed performance efficiency, cognitive implications, and responsible adoption in comparison with traditional and AI-assisted coding environments. Thirty participants, including professional developers and advanced computing students, completed equivalent programming tasks under three experimental conditions. Quantitative data were analyzed using descriptive statistics and repeated-measures ANOVA, while qualitative data were examined through thematic analysis. Results show that vibe coding significantly improved development efficiency, reducing task completion time by 27% compared with traditional coding and 12% compared with AI-assisted coding. However, these gains were accompanied by lower maintainability indices and higher security vulnerabilities, indicating trade-offs in software quality. Usability results yielded a good rating (SUS = 71.4), while cognitive workload remained moderate (NASA-TLX = 55.5), reflecting reduced syntactic effort but increased linguistic reasoning. Thematic analysis identified trust calibration, loss of control, cognitive adaptation, and prompt-engineering strategy as key constructs. Notably, perceived loss of control was associated with increased security risks due to reduced transparency and validation of AI-generated outputs. Based on these findings, the study proposes a three-pillar framework for responsible adoption: hybrid integration of human and AI capabilities, human oversight and transparent accountability, and context-aware deployment. Overall, vibe coding enhances productivity but requires critical oversight, reinforcing its role as a transformative yet transitional paradigm in software development.","authors":["Sales G. Aribe Jr.","Louie Jay S. Labastida"],"categories":["cs.SE","cs.AI","cs.ET"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09560","pdf_url":"https://arxiv.org/pdf/2609.09560","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助编程","人机交互","软件工程"],"reason":"研究AI辅助编程对开发者的影响，属于人机交互而非LLM仿真人类被试，无人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.09848","version":1,"title":"Subgroup Membership Inference Audits of Differentially Private Synthetic Text","zh_title":"差分隐私合成文本的子群体成员推断审计","abstract":"Synthetic data releases are increasingly proposed in the literature as a means of sharing realistic data replicas in lieu of sensitive private datasets. Even when the worst-case privacy leakage of such releases is bounded by means of differential privacy (DP), in practice a residual risk remains. Membership inference attack (MIA) audits are conducted to empirically quantify this risk. However, existing methods only measure average-case risk for randomly drawn records, which might conceal the risk to vulnerable subgroups. To highlight this issue, we define a subgroup-targeted membership inference game in which the target pool is an explicit parameter, and instantiate it with an audit of 32 proxies under three scenarios with different levels of attacker knowledge, across four datasets, three generators (DP-SGD fine-tuning, API-based prompting, and activation steering), and five privacy budgets. The audit shows that synthetic releases leak subgroup membership and that prior attacks systematically underestimate this leakage. DP is effective at the aggregate level: it substantially reduces average leakage at every budget we test. Three observations temper this picture. First, the remaining leakage is concentrated rather than spread out: under DP, a tenth of the records carries roughly 40% of it. Second, the protection DP delivers in practice is uneven: within its worst-case guarantee, the noise removes more of the measured leakage from random records than from high-risk ones---and a merged-pool audit that scores both record types against shared negatives confirms this at the record level. Third, \\emph{which} records leak proves to be a property of the release mechanism rather than of the record alone, so record-level risk cannot be assessed independently of the release.","authors":["Yidan Sun","Viktor Schlegel","Srinivasan Nandakumar","Siew Kei Lam","Anil Anthony Bharath"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09848","pdf_url":"https://arxiv.org/pdf/2609.09848","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["差分隐私","成员推断攻击","合成数据"],"reason":"研究差分隐私合成文本的成员推断审计，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.09331","version":1,"title":"Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities","zh_title":"短暂信息流与持久仪式：RushTok与基于事件的算法社区形成","abstract":"Each August, TikTok's For You page turns the University of Alabama's sorority recruitment into RushTok. We examine RushTok as an event-based algorithmic community: a collective assembled around a bounded offline ritual and sustained by recommendation. Using a mixed-methods survey (n=71) and a reflexive account of creator outreach, we ask who participates, how, and with what stakes. Findings show an ambiguous and entertainment based throughline; many called it a community (51/71) but few claimed membership (11/71). Affiliation centered on creators rather than shared practices, with parasocial attention clustering around a small set of potential new members (PNMs) and returning figures. Higher content exposure tracked with self-identification as a community member; those members commented, followed creators, and engaged across videos. Attempts to interview creators were met with silence or refusals, reflecting community boundary-work despite viral visibility. We outline implications for platform governance, including time-bounded context, graduated visibility, and aftercare.","authors":["Emelia Hughes","Tim Weninger"],"categories":["cs.HC","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09331","pdf_url":"https://arxiv.org/pdf/2609.09331","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["社交媒体","算法社区","平台治理"],"reason":"研究TikTok社区现象，不涉及LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.09333","version":1,"title":"Where Does the Human End? Creative Agency with Generative AI across Five Years of Chinese Digital Painting","zh_title":"人类止于何处？五年中国数字绘画中与生成式AI的创作能动性","abstract":"As generative AI enters creative work, practitioners must decide where AI assistance ends and human authorship begins. Human-agent interaction (HAI) research has examined AI as a tool, collaborator, consultant, and competitor. The longitudinal problem is how these roles are revised as systems become more capable, public, and economically embedded. We report a five-year interview study with 17 Chinese digital painters, based on annual semi-structured interviews from 2021 to 2025. Participants described recurring but non-uniform patterns of protective resistance, pragmatic task delegation, and, for some, reflective agency repartitioning. Early resistance protected observation, originality, signature, and ownership from AI. Later delegation placed AI in bounded tasks such as references, backgrounds, rough sketches, and client-facing drafts. By 2025, some participants built hybrid workflows around human-only zones, while others described fatigue, precarity, or difficulty locating a remaining human role. Peer norms, emotional climates, and production pressures shaped which delegations felt useful, acceptable, or exhausting. Copyright, authorship, and creative labor remained recurring limits on what participants were willing to delegate. We frame these accounts as longitudinal agency partitioning, the situated work of deciding which stages, responsibilities, values, and claims remain human in creative human-agent interaction. We discuss design implications for revisable agency-boundary controls, provenance scaffolds, and community-facing authorship norms.","authors":["Yibo Meng","Ruiqi Chen","Shuheng Cao","Weijia Zhang","Chengxi Zang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09333","pdf_url":"https://arxiv.org/pdf/2609.09333","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","创意能动性","生成式AI"],"reason":"研究人类与生成式AI在创作中的互动，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2609.09365","version":1,"title":"Echoes in the Algorithm: Analyzing the Fidelity of User Preferences Against Realized Platform Reach","zh_title":"算法中的回声：分析用户偏好与平台实际触达的保真度","abstract":"What does popular content look like when platforms withhold the usual cues? On TikTok, users still form impressions about which videos are taking off even when likes and view counts are hidden, delayed, or pushed to the margins of the interface. We study this problem through TokOrNot, a web-based game in which participants compared pairs of TikTok videos and reported (i) which one they preferred and (ii) which one they believed had reached a larger audience. We benchmark these judgments against verified public view counts, which we use as a bounded proxy for realized platform reach. Across 3,513 judgments from 363 participants, participants identified the higher-reach video only modestly above chance (56.75%, 95% CI: 56.01-58.55). Preference aligned with the higher-view video at a similar rate, while preference and prediction matched in 83.48% of trials (95% CI: 83.12-85.95). Performance also varied across content categories. Taken together, these results do not suggest that users can reliably read platform success from content alone. Instead, they point to a looser and more uncertain interpretive process in which reach judgments often track personal taste or other weak heuristics when explicit popularity cues are absent. We discuss the implications for algorithmic literacy and for interface designs that reduce visible metrics without leaving users to infer reach from uneven or idiosyncratic cues alone.","authors":["Emelia Hughes","Tim Weninger"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09365","pdf_url":"https://arxiv.org/pdf/2609.09365","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["用户行为","平台算法","人机交互"],"reason":"研究人类对TikTok视频的判断，不涉及LLM仿真人类被试，属于平台用户行为研…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.09713","version":1,"title":"How Far Do Capability Cues Travel? Anthropomorphism and Differentiated Trust in a Platform-Embedded AI Assistant","zh_title":"能力线索能传多远？平台嵌入式AI助手中的人形化与差异化信任","abstract":"Visible AI capabilities need not translate into broader judgments of trustworthiness. In a randomized 2 x 2 experiment with 270 U.S.-based Reddit users, an embedded assistant displayed one or three functions, with or without a brief rationale. Displaying three functions increased perceived multifunctionality; no other randomized main effect survived correction across the six outcomes. Rationale availability did not reliably increase perceived intelligence. Exploratory analysis indicated stronger uptake of the functional display at higher objective AI literacy. Among concurrently measured judgments, perceived multifunctionality was associated with perceived intelligence, which was associated with anthropomorphism and all three trust dimensions. After accounting for perceived intelligence, anthropomorphism was positively associated with benevolence, but not reliably with integrity or ability. These findings separate interface effects from relationships among users' perceptions and show why ability, integrity, and benevolence should be evaluated separately.","authors":["Chenchen Mao","Hanjing Shi","Haiyan Jia","Dominic DiFranzo"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09713","pdf_url":"https://arxiv.org/pdf/2609.09713","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","信任感知","AI助手"],"reason":"研究用户对AI助手的感知与信任，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.09321","version":1,"title":"Endorsement Without New Evidence: How Sequential Voting Inflates Mandates in Online Community Governance","zh_title":"无新证据的背书：在线社区治理中顺序投票如何夸大授权","abstract":"Online communities often treat large support margins in public elections as strong mandates. We argue that such margins can overstate the independent scrutiny behind a decision. Using 198,275 free-text rationales from Wikipedia admin elections, we introduce vote-text divergence, a measure that flags a decisive vote paired with a thin, deferential rationale. Divergence rises as voters arrive later, even after controlling for voter and election fixed effects. The pattern is consistent with information saturation: once prior text is accounted for, arrival order no longer predicts divergence, while accumulated prior evidence does. The effect is strongest among peripheral voters in the co-voting network. Yet divergence does not predict worse post-promotion outcomes, such as administrative activity or survival. Public tallies can therefore weaken the scrutiny signal even while selecting capable administrators: a margin may appear to reflect more consensus and support than it actually contains.","authors":["Zihan Chen","Lei Nico Zheng","Di Zhu"],"categories":["cs.SI","cs.CY","cs.HC"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09321","pdf_url":"https://arxiv.org/pdf/2609.09321","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["在线社区","投票行为","社会计算"],"reason":"研究在线社区投票行为，未使用LLM仿真人类被试，不涉及人类数据对照或LLM代理。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:27","error":null,"has_summary":false,"summary":null},{"id":"2609.09533","version":1,"title":"Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions","zh_title":"心理健康领域AI的可扩展监督：来自35万次治疗间AI辅导对话的经验","abstract":"Clinician review of every AI output is often proposed as a safeguard in mental healthcare, but vigilance research suggests this approach fails at scale and may paradoxically reduce safety. Drawing on our experience deploying an AI coaching tool across 350,000+ conversations between therapy sessions, we describe how we arrived at a three-layer human-on-the-loop oversight framework combining preventive design, real-time monitoring, and continuous clinician evaluation. We show how specific findings from clinical review drove iterative improvements, and offer practical recommendations for mental health professionals evaluating AI systems.","authors":["Matthew A. Scult","John L. Havlik","Kevin Ramotar","Ethan Goh","Manoj Kanagaraj"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09533","pdf_url":"https://arxiv.org/pdf/2609.09533","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI心理健康","人机对话","监督框架"],"reason":"AI心理健康辅导对话，属角色扮演聊天机器人，无实验或测量目的，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2609.10277","version":1,"title":"How neighbourhood ideology shapes misinformation belief in densely tied social networks","zh_title":"邻里意识形态如何影响紧密社交网络中的错误信息信念","abstract":"With the rapid spread of news on social media, understanding the propagation of misinformation is becoming increasingly important. One factor that affects individuals' vulnerability to false information is their ideological predisposition. Despite the large number of agent-based models that focus on social influence as a driver of the spread of false claims, they often fail to explicitly integrate personal ideological biases into belief formation. In this work, we explore how misinformation spreads through the interaction between individuals' ideological biases and social influence. Our model accounts for both the strength of individuals' ideological biases and the extent to which a false claim aligns with their ideology. Social influence modifies the effects of ideological intensity and false claim alignment through network interactions. Notably, the influence of neighbours' ideological intensity on belief is strongly affected by how well those neighbours are connected to one another. These results highlight the importance of considering both network structure and personal ideological biases when modelling misinformation propagation.","authors":["Soroush Karimi","Marcos Oliveira","Diogo Pacheco"],"categories":["cs.SI","cs.MA","physics.soc-ph"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10277","pdf_url":"https://arxiv.org/pdf/2609.10277","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["错误信息传播","基于代理的模型","社交网络"],"reason":"基于代理的模型模拟信息传播，未使用LLM，不涉及人类被试替代或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:37","error":null,"has_summary":false,"summary":null}]}