UX Research · Usability Testing · AI Platform
Foundry
Agent Builder
8 of 8 participants hit the same blocker, a usability catastrophe on Nielsen’s severity scale. The moderated study on Microsoft's agent-building platform turned a deprioritized known bug into a priority fix, and sent five findings to the Foundry design team.
Introduction
How might we make AI agent creation more intuitive for new developers?
Foundry set out to democratize agent creation, but new developers kept stalling before they finished.
Microsoft Foundry lets students, developers, and founders build AI agents. The vision is to democratize agent creation for people still learning, but a vision only matters if users can get through the flow.
On a Microsoft-sponsored project with the CoreAI team I set the research questions, designed the protocol, and ran all 8 moderated sessions. In session one, and every one after, we hit a blocker the team hadn't anticipated. The study went from “how intuitive is this?” to “this has to be fixed first.”
Note: Due to the confidential nature of this product and user privacy, visuals in this case study are limited.
Research Process
Research Questions
What we set out
to understand.
The question that framed everything: when does confusion stop being a learning curve and become a dead end?
Method
How we conducted
the study.
Eight moderated Think-Aloud sessions, chosen to catch not just where users failed but why.
I chose moderated testing because the research centered on mental models: unmoderated sessions capture task failure but not why. Think-Aloud exposed participants' reasoning in real time, and their silences were often the most revealing part.
Participant Criteria
I recruited participants matching Foundry's target audience: students and early-career professionals with technical backgrounds, curious about AI and looking to leverage it for projects.
What I tested, and what I left out
Eight sessions surface patterns, not statistics. I scoped to the first-run path and named what I left out, so when I reported that every participant hit the guardrail block the team could trust it as an unambiguous pattern rather than a stretched statistic.
Findings
What worked well.
Some patterns were already landing, and naming them tells the team what to protect.
Areas of Improvement
Five issues. One that stops everything.
One issue rated a Nielsen severity 4, a usability catastrophe, and it stopped all eight participants.
I ranked findings on Nielsen's severity scale. A confusing label is a different problem than a blocker stopping every user from finishing the core task. Severity framing gave the team a roadmap, not a list of complaints.
From findings to something the team could act on
A finding a team cannot act on is just an opinion. Every issue went to the team in the same shape: the observed behavior, how many of the eight it hit, a severity rating, a specific recommendation. The guardrail fix was flagged blocking, the terminology audit and landing-page hierarchy sequenced behind it.
Study at a Glance
The numbers behind
the research.
A study earns its keep by changing a decision. This one did.
What Changed
Two study outcomes trace back to what happened in the sessions.
A deprioritized bug became a priority fix. The guardrail block was already known internally and treated as minor. 8 of 8 participants failing the core task on it gave my mentor and the engineering team the evidence to re-rank it, moving a "minor annoyance" to a total blocker for first-run users.
Five findings entered the design team's queue. Each shipped with a severity rating and a concrete recommendation, with the terminology audit and landing-page hierarchy sequenced behind the guardrail fix.
"You dived into a complex product, asked all the right questions, and it's very clear that you put a lot of thought into planning and executing the study, and then translated that into a clear, engaging readout."
Research Mentor · Microsoft CoreAI
Reflection
What I learned & what surprised me.
The biggest lesson was adaptability. The guardrail blocker hit every participant from session one, forcing a real-time call: help them past it and lose data on its impact, or let them struggle and lose everything downstream. I chose a hybrid, letting them attempt it fully while I documented the confusion, then handing over a workaround so we could still test the rest.
When I flagged the guardrail issue, my mentor confirmed it was a known-but-underestimated bug, and my data gave engineering the evidence to prioritize the fix. Watching research directly flip a product decision was the highlight of this experience.
Thanks for reading
Let’s talk.
My one-page resume, or reach me directly.