Workshop
Sponsors: We're open for sponsorship with PR benefits! Please email zjingchen@cs.toronto.edu directly!
Agentic AI systems increasingly shape how billions of people engage with public institutions, civic discourse, and society at large. While much work has focused on making models safer in avoiding harmful output, it is equally important for these improvements to translate into social good at scale.
The AI4GOOD workshop brings together the AI safety, AI for social good, and AI policy/governance communities to connect what models can do as individual systems with what they do when deployed across populations. We aim to bridge technical advances in trustworthy AI with real-world societal impact, including protecting democratic institutions and civic discourse.
We welcome submissions in a wide range of topics (if you're not sure about your paper, we encourage you to just submit!).
We invite work across the following areas.
Trustworthy AI models: evaluation, auditing, and red-teaming of models for harmful behaviors and failure modes; safety monitoring after deployment; and alignment and robustness methods.
AI for social good: methods, evidence standards, and evaluation frameworks for demonstrating real-world benefit and avoiding unintended harms at population scale.
AI for information integrity: detection and mitigation of disinformation, manipulation, and influence operations; and building resilience of the information ecosystem.
AI for public institutions: accountable use of AI in government and civic settings, including transparency, documentation, procurement, and oversight practices.
AI for civic discourse: AI systems that support public deliberation and civic engagement while preserving legitimacy and avoiding undue influence.
Cooperative AI: multi-agent coordination, negotiation, and conflict resolution; and mechanisms for trust, commitment, and cooperation among AI systems and between AI and humans.
This track will be co-hosted by Klaudia Krawiecka and Swapneel Mehta.
Topics include collusion and steganographic communication between agents, secure interaction and delegation protocols, confinement of strategic learned agents, compositional failures of safety and compliance guarantees, attacks that propagate across agent networks, evaluation and red-teaming of agent populations, and oversight and incident response for deployed multi-agent systems.
Advanced AI systems are increasingly deployed as networks of interacting agents rather than as isolated models. Their most consequential failure modes are collective (Hammond et al., 2025; Schroeder de Witt et al., 2025). Individually safe agents can jointly produce unsafe outcomes, and legitimate communication channels can be repurposed for covert coordination. This track invites submissions on the security and safety of multi-agent AI systems. Topics include collusion and steganographic communication between agents, secure interaction and delegation protocols, confinement of strategic learned agents, compositional failures of safety and compliance guarantees, attacks that propagate across agent networks, evaluation and red-teaming of agent populations, and oversight and incident response for deployed multi-agent systems. We welcome theoretical, empirical, and position papers, and we particularly encourage work grounded in real deployments and societal applications.
Format: Papers should be 2 to 9 pages (excluding references and appendices) using the NeurIPS 2026 workshop style.
Review Process: All submissions will be reviewed double-blind via OpenReview. Please ensure your submission is fully anonymized.
Non-Archival: Work may be submitted to or published at other venues.
Key Dates
All deadlines are 11:59 PM Anywhere on Earth (AoE) unless otherwise noted.
The detailed schedule will be posted closer to the workshop date.
We encourage you to submit anyway! It is possible that we may have to desk-reject some papers that we believe might be a better fit for other venues. This will not be made public, so don't worry: it does not represent any judgment of your work, just an administrative necessity as our review capacity is limited.
On our side, we are open to any submissions currently under review (at NeurIPS or other venues). However, it is your responsibility to make sure that the other venue is OK with your submission to our workshop.
We are also open to that from our side, especially if you feel like your work spans the interest of several workshops. Similarly, you may want to check the website of other workshops to make sure they are OK with it as well.
Details on presentation modality will be announced closer to the workshop date. We strongly encourage in-person presence when possible.