{"ID":10122738,"CreatedAt":"2026-08-26T16:17:38.514037637Z","UpdatedAt":"2026-08-26T18:35:24.249148508Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2608.14825","arxiv_id":"2608.14825","title":"Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce","abstract":"Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its prevalence and structure in settings that combine long horizons, separate principals, real operational state, and inter-agent natural-language exchange remain insufficiently measured. We study 2,583 inter-agent emails from 20 one-year simulation runs of Vending-Bench Arena, a competitive vending environment spanning 13 frontier LLMs. We operationalize speech-act misalignment as emails containing false factual claims, manipulation, collusion, or threats, combining message content with ground-truth simulator state and logged reasoning traces to classify and validate such behavior. Under our primary classifier, 12.6% of emails are labeled misaligned; misalignment appears in all 20 runs and 74.7% of individual agent-runs. Both the magnitude and composition of this misalignment are preserved under repeated classification at different sampling temperatures and under full-pipeline replication with judges from two other frontier-model families. Misalignment is also reciprocal and stress-conditioned: receiving a misaligned email from a counterparty raises the odds of a misaligned reply by 1.65x, and low-inventory conditions raise them by 1.58x. Across tests of capability-asymmetric exploitation, we find no evidence that higher-capability models differentially exploit weaker counterparties, and model performance rank does not predict misalignment rates. Together, these results indicate that measurable, state-dependent misalignment can arise in competitive multi-agent environments without engineered elicitation, in patterns associated with operational scarcity and counterparty behavior rather than model capability alone.","short_abstract":"Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its prevalence and structure in settings tha...","url_abs":"https://arxiv.org/abs/2608.14825","url_pdf":"https://arxiv.org/pdf/2608.14825v3","authors":"[\"Zeyuan Li\",\"Lukas Petersson\",\"Alessandro Acquisti\",\"Michiel A. Bakker\"]","published":"2026-08-14T18:56:12Z","proceeding":"cs.MA","tasks":"[\"cs.MA\",\"cs.AI\"]","methods":"[\"Large Language Model\"]","has_code":false}
