Work Play About Resume

UX Research · Usability Testing · AI Platform

Foundry
Agent Builder

8 of 8 participants hit the same blocker, a usability catastrophe on Nielsen’s severity scale. The moderated study on Microsoft's agent-building platform turned a deprioritized known bug into a priority fix, and sent five findings to the Foundry design team.

My Role
UX Researcher
Team
4 UX researchers
Timeline
Jan – Mar 2026 · 10-week study
Engagement
Microsoft CoreAI · Sponsored Project
Methods
Moderated Usability Testing · Think Aloud
azure.microsoft.com/en-us/products/ai-foundry
Azure AI Foundry website
Explore the live platform ↗
At a glance
01
The challenge
A bug the team had logged and deprioritized was quietly blocking first-run users from creating an agent on Microsoft Foundry. On Nielsen’s severity scale it rates a 4, a usability catastrophe that prevents task completion.
02
What I did
On a team of four researchers, I ran all 8 moderated think-aloud sessions on the live platform, framed every issue by severity, and handed a prioritized set to the Foundry design and engineering teams.
03
The outcome
Five findings entered the design team's queue, each with a severity rating and a concrete recommendation.

Introduction

How might we make AI agent creation more intuitive for new developers?

Foundry set out to democratize agent creation, but new developers kept stalling before they finished.

Microsoft Foundry lets students, developers, and founders build AI agents. The vision is to democratize agent creation for people still learning, but a vision only matters if users can get through the flow.

On a Microsoft-sponsored project with the CoreAI team I set the research questions, designed the protocol, and ran all 8 moderated sessions. In session one, and every one after, we hit a blocker the team hadn't anticipated. The study went from “how intuitive is this?” to “this has to be fixed first.”

Note: Due to the confidential nature of this product and user privacy, visuals in this case study are limited.

Research Process

Scoping
Defined research questions with the design team
Protocol Design
Designed study protocol and task scenarios
8 Sessions
Moderated think-aloud usability sessions
Synthesis
Coded findings: 5 findings, including 1 critical blocker
Readout
Presented prioritized recommendations to CoreAI team

Research Questions

What we set out
to understand.

The question that framed everything: when does confusion stop being a learning curve and become a dead end?

01
At what point in the Foundry agent-creation flow do users first feel confused or overwhelmed?
02
Which concepts or terms make users feel like they need external help to continue?
03
Do users perceive confusion as a temporary learning curve or a hard blocker that makes the tool unusable?
04
What data types do users want to work with, and what do they want their agents to do?

Method

How we conducted
the study.

Eight moderated Think-Aloud sessions, chosen to catch not just where users failed but why.

I chose moderated testing because the research centered on mental models: unmoderated sessions capture task failure but not why. Think-Aloud exposed participants' reasoning in real time, and their silences were often the most revealing part.

Moderated Sessions
One-on-one moderated usability tests with real tasks in the live Foundry platform, using Think Aloud protocol to capture real-time reasoning and confusion points.
n = 8 participants
Task-Based Scenarios
Participants completed structured tasks including agent creation, configuring knowledge sources, uploading data, and publishing agents to mirror real use cases.
Single condition, all 8
Likert Ratings & Quotes
Post-task ease-of-use ratings on a 5-point scale where 1 is very easy and 5 is very difficult, combined with qualitative Think Aloud data and direct participant quotes for triangulated findings.
Mixed methods

Participant Criteria

I recruited participants matching Foundry's target audience: students and early-career professionals with technical backgrounds, curious about AI and looking to leverage it for projects.

Critical Criteria
Active student status CS or technical background Curious about AI Looking to leverage AI for a project Experience with data sets

What I tested, and what I left out

Eight sessions surface patterns, not statistics. I scoped to the first-run path and named what I left out, so when I reported that every participant hit the guardrail block the team could trust it as an unambiguous pattern rather than a stretched statistic.

In Scope
+First-run agent creation, from empty state to a published agent
+The Knowledge, Data, and FoundryIQ features, on the live platform
+Where confusion turns into a hard stop, framed by severity
Deliberately Out of Scope
~End-to-end integration of a published agent into a real product, a study of its own
~Statistical significance: with eight participants I report patterns, not population percentages
~Internal telemetry and usage analytics, which sat outside my access as an external researcher

Findings

What worked well.

Some patterns were already landing, and naming them tells the team what to protect.

Publishing an Agent
5 out of 8 participants rated publishing their agent very easy. The publishing flow aligned with familiar patterns from other tools, making it intuitive and frictionless.
"Publishing the agent was pretty straightforward and aligns with what I would expect it to do because it's very similar UI elements to what other tools do right now."
Overall Task Completion
4 out of 8 participants rated overall ease of use a 2 on the 1 to 5 scale, one step below very easy. While the platform had areas of confusion, participants were ultimately able to navigate and complete tasks, suggesting a solid foundation to build on.

Areas of Improvement

Five issues. One that stops everything.

One issue rated a Nielsen severity 4, a usability catastrophe, and it stopped all eight participants.

I ranked findings on Nielsen's severity scale. A confusing label is a different problem than a blocker stopping every user from finishing the core task. Severity framing gave the team a roadmap, not a list of complaints.

Severity Scale (Nielsen's)
4 Usability catastrophe: Prevents task completion; imperative to fix
3 Major problem: Causes significant confusion; important to fix
2 Minor problem: Adds friction but doesn't block completion; should be addressed
1 Cosmetic: Surface-level issue; fix if time permits
Severity 4 Guardrail Blocks Agent Creation
8 out of 8 participants hit a hard stop creating an agent because the interface blocked interactions due to an unassigned or mismanaged guardrail, so the agent could not run until one was set. None got past it unaided; I documented the confusion in full, then handed each participant a workaround so the rest of the flow could still be tested, which is why later tasks show completions. The system leaves new agents in an ambiguous "inheriting" state instead of automatically assigning Microsoft's default guardrail, causing the interface to prevent users from completing the task.
The error message further compounds confusion by prompting users to "Create Guardrail," when the actual resolution is to reassign to an existing default guardrail. Participants lacked contextual guidance on why the interaction was blocked, were unclear about the differences between guardrail versions, and could not see the active guardrail status, all of which increased frustration and wasted time.
Recommendations
Automatically assign Microsoft's default guardrail during agent creation. Update error messages to direct users to "Reassign Guardrail" rather than "Create Guardrail." Provide inline explanations about why interactions are blocked and how to resolve them. Clarify guardrail purposes and version differences through tooltips or descriptions. Make guardrail status more visible in the interface.
Severity 3 Confusing Terminology & Labeling
7 out of 8 participants were confused by overlapping or unclear terminology in the platform. Key points of confusion included the distinction between "Tools" and "Knowledge" (and why file uploads appeared under Tools), as well as the difference between "Agent Instructions" and "Message Agent." While this didn't fully prevent task completion, it was the most significant usability friction in the overall experience.
Recommendations
Audit and simplify terminology across the platform to ensure labels are distinct, descriptive, and consistent. Add contextual definitions (tooltips or inline descriptions) for key concepts like "Tools," "Knowledge," and "Agent Instructions." Consider renaming overlapping terms to reduce cognitive load for first-time users.
Severity 2 Cluttered Landing Page & Visual Hierarchy
4 out of 8 participants experienced discoverability issues on the landing page. The "Start Building" button lacked visual prominence due to its relatively small size and the presence of multiple competing visual elements. The "Coding Quick Start" bar was significantly larger and attracted users' attention first. Additionally, the similarity in terminology between these two options created confusion regarding the appropriate starting point.
"I see a couple of places I could go to. Do I go to Start Building? Do I go to the Coding Quick Start part?"
Recommendations
Rebalance the page layout to reduce visual competition among elements and enhance hierarchical clarity. Increase the visual prominence of the "Start Building" button by enlarging its size, contrast, and positioning it within a primary focal area.
Severity 2 Unclear Platform Navigation Flow
Each of the 8 participants navigated Microsoft Foundry via a different flow. After creating an agent, 2 out of 8 participants interacted with the navigation sidebar tabs to gain more understanding about the platform's terminology. While participants were ultimately able to navigate, the lack of guided structure forced exploratory behavior.
"I'll start by looking at the nav bar because there was no clear instruction of how like the different steps that will be involved in making the AI agents, so I'll have to explore the software on my own."
Recommendations
Introduce onboarding assistance or a guided tutorial for first-time users. 3 out of 8 participants specifically suggested this. As one shared: "I feel like if I looked up a tutorial or if the platform gave me some info when I created my account, it would be pretty easy to figure out as you go."
Severity 2 Misleading Error Messages on File Upload
5 out of 8 participants experienced confusion when files were successfully uploaded but an error message appeared. The upload process itself was straightforward, but the false error introduced unnecessary doubt and broke user confidence in the system.
"At least uploading [files] was straightforward. But that error message was a little bit confusing."
Recommendations
Investigate and resolve the underlying bug causing false error messages on successful uploads. Ensure confirmation states clearly communicate success and distinguish between warnings, errors, and informational messages.

From findings to something the team could act on

A finding a team cannot act on is just an opinion. Every issue went to the team in the same shape: the observed behavior, how many of the eight it hit, a severity rating, a specific recommendation. The guardrail fix was flagged blocking, the terminology audit and landing-page hierarchy sequenced behind it.

How one finding became a priority fix
Observed
8 of 8 users left a new agent in an ambiguous "inheriting" guardrail state and hit a hard stop.
Severity
Rated Nielsen severity 4, a usability catastrophe: a total blocker for first-run users, not a minor annoyance.
Recommendation
Assign a default guardrail, and rewrite the error to point to the real fix.
What changed
A deprioritized bug was re-ranked to a priority fix; five findings entered the design queue.

Study at a Glance

The numbers behind
the research.

A study earns its keep by changing a decision. This one did.

8/8
participants hit the same critical guardrail blocker
5
distinct usability issues identified and prioritized by severity for the Foundry team
2
study outcomes: a deprioritized bug re-ranked to a priority fix, and five findings sequenced into the design queue

What Changed

Two study outcomes trace back to what happened in the sessions.

A deprioritized bug became a priority fix. The guardrail block was already known internally and treated as minor. 8 of 8 participants failing the core task on it gave my mentor and the engineering team the evidence to re-rank it, moving a "minor annoyance" to a total blocker for first-run users.

Five findings entered the design team's queue. Each shipped with a severity rating and a concrete recommendation, with the terminology audit and landing-page hierarchy sequenced behind the guardrail fix.

"You dived into a complex product, asked all the right questions, and it's very clear that you put a lot of thought into planning and executing the study, and then translated that into a clear, engaging readout."

Research Mentor · Microsoft CoreAI

Reflection

What I learned & what surprised me.

The biggest lesson was adaptability. The guardrail blocker hit every participant from session one, forcing a real-time call: help them past it and lose data on its impact, or let them struggle and lose everything downstream. I chose a hybrid, letting them attempt it fully while I documented the confusion, then handing over a workaround so we could still test the rest.

When I flagged the guardrail issue, my mentor confirmed it was a known-but-underestimated bug, and my data gave engineering the evidence to prioritize the fix. Watching research directly flip a product decision was the highlight of this experience.

What Went Well
+ Timely and sufficient participant recruitment: I met the target of 6–8 participants
+ Biweekly 1-hour check-ins with my Microsoft mentor helped me resolve issues in real time
+ Strong team organization with weekly internal syncs
+ Each usability session was productive and surfaced new insights
What I'd Do Differently
~ Run a pilot session before the formal study: a dry run would have surfaced the guardrail blocker earlier and given me time to design a cleaner workaround protocol
~ Include a broader participant pool: startup founders (not just students) would have revealed whether the terminology issues are universal or expertise-dependent
~ Add a retrospective interview after each session: some of my best insights came from off-script comments, and a structured debrief would have captured more of them

Thanks for reading

Let’s talk.

My one-page resume, or reach me directly.

Next Project
Foundry: Models Redesign
View case study →