The Treacherous Envoy Problem: Trust, Collusion, and Accountability in Multi-Agent Workflows

SACMAT |

Organized by ACM

As LLM-based agents shift from generating text to executing real workflows, they negotiate across organizational boundaries where principals have conflict of interests, invoke external tools, and trigger irreversible side effects. This shift moves humans from doers to verifiers: the central security challenge becomes validating that an agent’s actions match the principal’s intent. We argue that this validation and verification (V&V) challenge is subject to a structural trilemma: strengthening any one of (1) expressive natural-language negotiation, (2) verifiable conformance of effects to intent, and (3) bounded disclosure that protects principals’ private context tightens the constraints on the other two. We formalize this as the Treacherous Envoy Problem (TEP), identify five recurring tensions where the trilemma creates practical conflicts, categorize five exploitation patterns that adversaries use to leverage it, and provide formal evidence that the three properties resist independent resolution. We conclude with research directions toward building the V&V infrastructure that agent workflows currently lack.