UX Research · Usability Testing · AI Platform
Foundry
Agent Builder
8 of 8 participants hit the same Severity-4 blocker, turning a deprioritized known bug into a priority fix and sending five findings to the Foundry design team. This is the moderated usability study on Microsoft's AI agent-building platform that surfaced it.
Introduction
How might we make AI agent creation more intuitive for new developers?
Foundry set out to democratize agent creation, but new developers kept stalling before they finished.
Microsoft Foundry lets students, developers, and startup founders build AI agents. The vision is to democratize agent creation for people still learning, but vision only matters if users can get through the flow.
As a UX Researcher on Microsoft's CoreAI team, I led 8 moderated usability sessions on the Knowledge, Data, and FoundryIQ features: I defined the research questions, designed the protocol, ran the sessions, and owned team communication. In the very first session, and every one after, we found a critical blocker that 100% of participants hit, one the team hadn't anticipated. The study went from "let's see how intuitive the flow is" to "we need to fix this before anything else matters."
Note: Due to the confidential nature of this product and user privacy, visuals in this case study are limited.
Research Process
Research Questions
What we set out
to understand.
The question that framed everything: when does confusion stop being a learning curve and become a dead end?
I shaped the questions to go past surface usability and find where confusion turns from "learning curve" into "I'm done." That distinction, temporary confusion versus a hard blocker, became the most important framing of the study.
Method
How we conducted
the study.
Eight moderated Think-Aloud sessions, chosen to catch not just where users failed but why.
I chose moderated usability testing because the research centered on mental models. I needed to be there when confusion happened. Remote unmoderated testing would capture task failure, but not why. The Think Aloud protocol revealed participants' reasoning in real time, and their silences were often most revealing.
Participant Criteria
I recruited participants matching Foundry's target audience: students and early-career professionals with technical backgrounds, curious about AI and looking to leverage it for projects.
What I tested, and what I left out
Eight sessions surface pattern-level problems, not statistics. I scoped the study to the first-run agent-creation path (create, add knowledge and data, configure, publish) and left three things out of frame on purpose.
Naming these limits up front kept the readout honest: when I reported that every participant hit the guardrail block, the team could trust it as a qualitative certainty, not a statistic stretched past what eight sessions can carry.
Findings
What worked well.
Some patterns were already landing, and naming them tells the team what to protect.
Not everything was broken, and that matters. Naming what works tells the team which patterns to protect as they iterate.
Areas of Improvement
Five issues. One that stops everything.
One issue was a Severity-4 blocker that stopped all eight participants; the rest, I ranked behind it.
I prioritized findings using Nielsen's severity scale. A confusing label is a different problem than a blocker preventing every user from completing the core task. Clear severity framing gave the Foundry team an actionable roadmap, not just a list of complaints.
From findings to something the team could act on
A finding a team cannot act on is just an opinion. I structured every issue the same way, so design and engineering could pick up their part without a translation step: the observed behavior, how many of the eight it hit, a severity rating, and a specific recommendation. I handed the set to the Foundry design team with the guardrail fix flagged as blocking, and the terminology audit and landing-page hierarchy sequenced behind it for later sprints.
Study at a Glance
The numbers behind
the research.
One study, two decisions changed: a deprioritized bug became the team's top fix.
Eight sessions. Five distinct findings. One critical blocker affected every single participant, a finding so consistent it immediately became the team's top priority.
What Changed
A study earns its keep by changing a decision. This one changed two.
A deprioritized bug became a priority fix. The guardrail block was already known internally and treated as minor. 8 of 8 participants failing the core task on it gave my mentor and the engineering team the evidence to re-rank it, moving a "minor annoyance" to a total blocker for first-run users. Watching that call flip on the strength of research was the clearest proof of impact I saw.
Five findings entered the design team's queue. Each shipped with a severity rating and a concrete recommendation, so the team got a prioritized roadmap rather than a list of complaints, with the terminology audit and landing-page hierarchy sequenced behind the guardrail fix.
"You dived into a complex product, asked all the right questions, and it's very clear that you put a lot of thought into planning and executing the study, and then translated that into a clear, engaging readout."
Research Mentor · Microsoft CoreAI
Reflection
What I learned & what surprised me.
The biggest lesson was adaptability. The guardrail blocker hit every participant from Session 1, forcing a real-time call: help them past it and lose data on the blocker's impact, or let them struggle and lose downstream data. I chose a hybrid: participants attempted the task fully while I documented the confusion, then I gave them a workaround so we could still test the rest of the flow. That preserved both the critical finding and the downstream insights.
When I flagged the guardrail issue, my mentor confirmed it was a known-but-underestimated bug, and my data gave engineering the evidence to prioritize the fix. Watching research directly flip a product decision was the highlight of this experience.