Research

I study how multi-agent AI systems fail and how to make them reliable, alongside applied machine-learning work. AgentCollabBench is the published piece of that; further work is in progress and will appear here when there is evidence to share.

2026

Publication · Preprint

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators

A diagnostic benchmark of 900 human-validated tasks that exposes how multi-agent systems silently corrupt reasoning chains even when the final output looks correct.

multi-agent systems · LLM evaluation