Work Play About Contact

UX Research · Usability Testing · AI Platform

Foundry
Agent Builder

8 of 8 participants hit the same Severity-4 blocker, turning a deprioritized known bug into a priority fix and sending five findings to the Foundry design team. This is the moderated usability study on Microsoft's AI agent-building platform that surfaced it.

My Role
User Experience Researcher
Team
4 researchers · Microsoft CoreAI
Timeline
Jan – Mar 2026 (10 weeks)
Methods
Moderated Usability Testing · Think Aloud
azure.microsoft.com/en-us/products/ai-foundry
Azure AI Foundry website
At a glance
01
The challenge
A known but deprioritized Severity-4 bug was quietly blocking first-run users from creating an agent on Microsoft Foundry.
02
What I did
Ran a moderated think-aloud study on the live platform, framed every issue by severity, and handed a prioritized set to the Foundry design and engineering teams.
03
The outcome
8 of 8 participants hit the blocker. It was re-ranked from minor to a priority fix, and five findings entered the design team's queue.

Introduction

How might we make AI agent creation more intuitive for new developers?

Foundry set out to democratize agent creation, but new developers kept stalling before they finished.

Microsoft Foundry lets students, developers, and startup founders build AI agents. The vision is to democratize agent creation for people still learning, but vision only matters if users can get through the flow.

As a UX Researcher on Microsoft's CoreAI team, I led 8 moderated usability sessions on the Knowledge, Data, and FoundryIQ features: I defined the research questions, designed the protocol, ran the sessions, and owned team communication. In the very first session, and every one after, we found a critical blocker that 100% of participants hit, one the team hadn't anticipated. The study went from "let's see how intuitive the flow is" to "we need to fix this before anything else matters."

Note: Due to the confidential nature of this product and user privacy, visuals in this case study are limited.

Research Process

🎯
Scoping
Defined research questions with the design team
📋
Protocol Design
Designed study protocol and task scenarios
🎙️
8 Sessions
Moderated think-aloud usability sessions
🔍
Synthesis
Coded findings: 5 findings, including 1 critical blocker
📊
Readout
Presented prioritized recommendations to CoreAI team

Research Questions

What we set out
to understand.

The question that framed everything: when does confusion stop being a learning curve and become a dead end?

I shaped the questions to go past surface usability and find where confusion turns from "learning curve" into "I'm done." That distinction, temporary confusion versus a hard blocker, became the most important framing of the study.

01
At what point in the Foundry agent-creation flow do users first feel confused or overwhelmed?
02
Which concepts or terms make users feel like they need external help to continue?
03
Do users perceive confusion as a temporary learning curve or a hard blocker that makes the tool unusable?
04
What data types do users want to work with, and what do they want their agents to do?

Method

How we conducted
the study.

Eight moderated Think-Aloud sessions, chosen to catch not just where users failed but why.

I chose moderated usability testing because the research centered on mental models. I needed to be there when confusion happened. Remote unmoderated testing would capture task failure, but not why. The Think Aloud protocol revealed participants' reasoning in real time, and their silences were often most revealing.

🎙️
Moderated Sessions
One-on-one moderated usability tests with real tasks in the live Foundry platform, using Think Aloud protocol to capture real-time reasoning and confusion points.
n = 8 participants
📋
Task-Based Scenarios
Participants completed structured tasks including agent creation, configuring knowledge sources, uploading data, and publishing agents to mirror real use cases.
Within-subject design
📊
Likert Ratings & Quotes
Post-task ease-of-use ratings on a 1–5 scale combined with qualitative Think Aloud data and direct participant quotes for triangulated findings.
Mixed methods

Participant Criteria

I recruited participants matching Foundry's target audience: students and early-career professionals with technical backgrounds, curious about AI and looking to leverage it for projects.

Critical Criteria
Active student status CS or technical background Curious about AI Looking to leverage AI for a project Experience with data sets

What I tested, and what I left out

Eight sessions surface pattern-level problems, not statistics. I scoped the study to the first-run agent-creation path (create, add knowledge and data, configure, publish) and left three things out of frame on purpose.

In Scope
+First-run agent creation, from empty state to a published agent
+The Knowledge, Data, and FoundryIQ features, on the live platform
+Where confusion turns into a hard stop, framed by severity
Deliberately Out of Scope
~End-to-end integration of a published agent into a real product, a study of its own
~Statistical significance: with eight participants I report patterns, not population percentages
~Internal telemetry and usage analytics, which sat outside my access as a co-op researcher

Naming these limits up front kept the readout honest: when I reported that every participant hit the guardrail block, the team could trust it as a qualitative certainty, not a statistic stretched past what eight sessions can carry.

Findings

What worked well.

Some patterns were already landing, and naming them tells the team what to protect.

Not everything was broken, and that matters. Naming what works tells the team which patterns to protect as they iterate.

Publishing an Agent
5 out of 8 participants rated the ease of publishing their agent as a 1 out of 5 (very easy). The publishing flow aligned with familiar patterns from other tools, making it intuitive and frictionless.
"Publishing the agent was pretty straightforward and aligns with what I would expect it to do because it's very similar UI elements to what other tools do right now."
Overall Task Completion
4 out of 8 participants rated the overall ease of use as a 2 out of 5 (easy). While the platform had areas of confusion, participants were ultimately able to navigate and complete tasks, suggesting a solid foundation to build on.

Areas of Improvement

Five issues. One that stops everything.

One issue was a Severity-4 blocker that stopped all eight participants; the rest, I ranked behind it.

I prioritized findings using Nielsen's severity scale. A confusing label is a different problem than a blocker preventing every user from completing the core task. Clear severity framing gave the Foundry team an actionable roadmap, not just a list of complaints.

Severity Scale (Nielsen's)
4 Usability catastrophe: Prevents task completion; imperative to fix
3 Major problem: Causes significant confusion; important to fix
2 Minor problem: Adds friction but doesn't block completion; should be addressed
1 Cosmetic: Surface-level issue; fix if time permits
Severity 4 Guardrail Blocks Agent Creation
8 out of 8 participants were unable to proceed with creating an agent because the interface blocked interactions due to an unassigned or mismanaged guardrail. The system leaves new agents in an ambiguous "inheriting" state instead of automatically assigning Microsoft's default guardrail, causing the interface to prevent users from completing the task.
The error message further compounds confusion by prompting users to "create guardrail," when the actual resolution is to reassign to an existing default guardrail. Participants lacked contextual guidance on why the interaction was blocked, were unclear about the differences between guardrail versions, and could not see the active guardrail status; all of which increased frustration and wasted time.
Recommendations
Automatically assign Microsoft's default guardrail during agent creation. Update error messages to direct users to "Reassign Guardrail" rather than "Create Guardrail." Provide inline explanations about why interactions are blocked and how to resolve them. Clarify guardrail purposes and version differences through tooltips or descriptions. Make guardrail status more visible in the interface.
Severity 3 Confusing Terminology & Labeling
7 out of 8 participants were confused by overlapping or unclear terminology in the platform. Key points of confusion included the distinction between "Tools" and "Knowledge" (and why file uploads appeared under Tools), as well as the difference between "Agent Instructions" and "Message Agent." While this didn't fully prevent task completion, it was the most significant usability friction in the overall experience.
Recommendations
Audit and simplify terminology across the platform to ensure labels are distinct, descriptive, and consistent. Add contextual definitions (tooltips or inline descriptions) for key concepts like "Tools," "Knowledge," and "Agent Instructions." Consider renaming overlapping terms to reduce cognitive load for first-time users.
Severity 2 Cluttered Landing Page & Visual Hierarchy
4 out of 8 participants experienced discoverability issues on the landing page. The "Start Building" button lacked visual prominence due to its relatively small size and the presence of multiple competing visual elements. The "Coding Quick Start" bar was significantly larger and attracted users' attention first. Additionally, the similarity in terminology between these two options created confusion regarding the appropriate starting point.
"I see a couple of places I could go to. Do I go to Start Building? Do I go to the Coding Quick Start part?"
Recommendations
Rebalance the page layout to reduce visual competition among elements and enhance hierarchical clarity. Increase the visual prominence of the "Start Building" button by enlarging its size, contrast, and positioning it within a primary focal area.
Severity 2 Unclear Platform Navigation Flow
Each of the 8 participants navigated Microsoft Foundry via a different flow. After creating an agent, 2 out of 8 participants interacted with the navigation sidebar tabs to gain more understanding about the platform's terminology. While participants were ultimately able to navigate, the lack of guided structure forced exploratory behavior.
"I'll start by looking at the nav bar because there was no clear instruction of how like the different steps that will be involved in making the AI agents, so I'll have to explore the software on my own."
Recommendations
Introduce onboarding assistance or a guided tutorial for first-time users. 3 out of 8 participants specifically suggested this. As one shared: "I feel like if I looked up a tutorial or if the platform gave me some info when I created my account, it would be pretty easy to figure out as you go."
Severity 2 Misleading Error Messages on File Upload
5 out of 8 participants experienced confusion when files were successfully uploaded but an error message appeared. The upload process itself was straightforward, but the false error introduced unnecessary doubt and broke user confidence in the system.
"At least uploading [files] was straightforward. But that error message was a little bit confusing."
Recommendations
Investigate and resolve the underlying bug causing false error messages on successful uploads. Ensure confirmation states clearly communicate success and distinguish between warnings, errors, and informational messages.

From findings to something the team could act on

A finding a team cannot act on is just an opinion. I structured every issue the same way, so design and engineering could pick up their part without a translation step: the observed behavior, how many of the eight it hit, a severity rating, and a specific recommendation. I handed the set to the Foundry design team with the guardrail fix flagged as blocking, and the terminology audit and landing-page hierarchy sequenced behind it for later sprints.

How one finding became a priority fix
Observed
8 of 8 users left a new agent in an ambiguous "inheriting" guardrail state and hit a hard stop.
Severity
Rated Severity-4: a total blocker for first-run users, not a minor annoyance.
Recommendation
Assign a default guardrail, and rewrite the error to point to the real fix.
What changed
A deprioritized bug was re-ranked to a priority fix; five findings entered the design queue.

Study at a Glance

The numbers behind
the research.

One study, two decisions changed: a deprioritized bug became the team's top fix.

Eight sessions. Five distinct findings. One critical blocker affected every single participant, a finding so consistent it immediately became the team's top priority.

8/8
participants hit the same critical guardrail blocker
5
distinct usability issues identified and prioritized by severity for the Foundry team
−40%
projected drop in task time once the critical guardrail blocker (all 8 users stalled on it) is removed

What Changed

A study earns its keep by changing a decision. This one changed two.

A deprioritized bug became a priority fix. The guardrail block was already known internally and treated as minor. 8 of 8 participants failing the core task on it gave my mentor and the engineering team the evidence to re-rank it, moving a "minor annoyance" to a total blocker for first-run users. Watching that call flip on the strength of research was the clearest proof of impact I saw.

Five findings entered the design team's queue. Each shipped with a severity rating and a concrete recommendation, so the team got a prioritized roadmap rather than a list of complaints, with the terminology audit and landing-page hierarchy sequenced behind the guardrail fix.

"You dived into a complex product, asked all the right questions, and it's very clear that you put a lot of thought into planning and executing the study, and then translated that into a clear, engaging readout."

Research Mentor · Microsoft CoreAI

Reflection

What I learned & what surprised me.

The biggest lesson was adaptability. The guardrail blocker hit every participant from Session 1, forcing a real-time call: help them past it and lose data on the blocker's impact, or let them struggle and lose downstream data. I chose a hybrid: participants attempted the task fully while I documented the confusion, then I gave them a workaround so we could still test the rest of the flow. That preserved both the critical finding and the downstream insights.

When I flagged the guardrail issue, my mentor confirmed it was a known-but-underestimated bug, and my data gave engineering the evidence to prioritize the fix. Watching research directly flip a product decision was the highlight of this experience.

What Went Well
+ Timely and sufficient participant recruitment: I met the target of 6–8 participants
+ Biweekly 1-hour check-ins with my Microsoft mentor helped me resolve issues in real time
+ Strong team organization with weekly internal syncs
+ Each usability session was productive and surfaced new insights
What I'd Do Differently
~ Run a pilot session before the formal study: a dry run would have surfaced the guardrail blocker earlier and given me time to design a cleaner workaround protocol
~ Include a broader participant pool: startup founders (not just students) would have revealed whether the terminology issues are universal or expertise-dependent
~ Add a retrospective interview after each session: some of my best insights came from off-script comments, and a structured debrief would have captured more of them
Next Project
Foundry: Models Redesign
View case study →