Skip to content

Research

I research how to make AI agents safe and reliable enough to trust with real software engineering work.

I'm a PhD student in Software Engineering at Polytechnique Montréal. I work on AI agents: how they fail, and what it would take to trust them with real engineering work. The same question pulls me into AI for software engineering (AI4SE).

Trustworthy AI Agents

How do we make LLM-based agents safe and reliable enough to act autonomously?

Agents are being handed real responsibilities faster than anyone can characterize how they fail. A chatbot's mistake is a bad answer; an agent's mistake is an action, and it lands somewhere. My PhD work at Polytechnique Montréal builds evaluations and guardrails for agentic systems, starting from their observed failure modes.

AI for Software Engineering

Can LLM ensembles catch what single models miss in software artifacts?

Incomplete requirements are a quiet, expensive failure mode: what a specification leaves unsaid rarely surfaces until it breaks something downstream. I mapped how the field defines, measures, and tools requirements completeness, then showed that ensembles of diverse LLMs catch and repair gaps single models miss. Next I want to try the same ensemble approach on other software artifacts.

Compute-Optimal LLM Reasoning

When is more inference compute actually worth it?

Spending more compute at inference time can substitute for model scale, but the trade-off is rarely measured carefully. With Pavly Halim, I compared inference-scaling strategies against reasoning-trained approaches on math problem-solving under matched compute budgets. The question gets more relevant every year: for deployed systems, inference is where the money goes.

Interested in collaborating?

I'm always glad to talk about AI agents, AI safety & reliability, and AI for software engineering (AI4SE). I also take on advising roles — an outside eye on agent systems and AI4SE problems.

Get in touch