Research
I study how multi-agent AI systems fail and how to make them reliable, alongside applied machine-learning work. AgentCollabBench is the published piece of that; further work is in progress and will appear here when there is evidence to share.
2026
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
A diagnostic benchmark of 900 human-validated tasks that exposes how multi-agent systems silently corrupt reasoning chains even when the final output looks correct.
multi-agent systems · LLM evaluation