Trust Calibration case study cover

Trust Calibration for Agentic AI in Enterprise

Agentic AI products are designed for users who already trust them. Most don't. This is what designing for the gap looks like.

Role
Product Manager and Designer
Timeline
2 weeks
Tools
Claude Code, Figma Make
Problem Statement

The problem nobody has designed for

Agentic AI tools are being deployed into organizations where most users have no prior experience working with an AI agent. These users arrive without a mental model of what the agent does, how it reasons, or what happens when it gets something wrong.

The products themselves assume a level of familiarity the user hasn't earned yet. They ask for action-level trust before establishing baseline trust. The result is low adoption, shallow usage, and AI investment that doesn't convert into measurable business value.

This isn't a UI problem. It's a sequencing problem. The industry hasn't designed for the moment before the agent can be useful.

Diagnosis

Layers of trust, but the base is missing.

Baseline trust is the belief that a system is predictable and understandable - that you know roughly what it will do and why. Action-level trust is the willingness to act on a system's output and put your name on the result. These are different things, and they have to be earned in sequence. You can't skip baseline trust and expect action-level trust to follow.

Agentic AI products are designed around action-level trust. They surface recommendations, flag anomalies, and generate outputs. What they rarely do is show their work in a way that builds the baseline layer first. The assumption is that users will calibrate through experience. But in organizational contexts, most users don't have the time, the safety, or the permission to experiment before they're expected to produce.

Even reasoning traces don't close the gap. A trace answers "what did it do" after the fact. Calibration answers "can I rely on it" before the first real case. Traces are the raw material for trust - calibration is the mechanism that builds it.

Trust pyramid diagram showing baseline trust as the foundation for action-level trust
Adapted from the NN/g Pyramid of Trust
User Profile

Use case

Maya is a mid-level lending auditor at an equipment rental company. As part of a firm-wide AI initiative, her team was given an enterprise AI workspace built for regulated industries - an AI assistant, search across the company's internal systems, and customizable agents that automate multi-step workflows.

On paper, it should transform her work. A lending audit means pulling a file - application, bank statements, audit reports, approval memos - and verifying the decision followed policy. The workspace can search and summarize documents scattered across the company's systems, extract key figures, flag inconsistencies between what was documented and what policy requires, and draft the audit summary. A day of manual document review should become a couple of hours of reviewing the agent's work.

But Maya's confusion sits in her workflow, not the product. The workspace logs its interactions and shows reasoning traces - the tool itself is well built. What nobody told her is how to work with it. The capability was deployed. The working relationship wasn't.

Delegation

Which parts of the audit is she allowed to delegate to the agent, and which must she still do herself?

Defensibility

Is an agent-drafted summary defensible when a regulator asks who verified the bank statements?

Judgment

How does the agent decide a discrepancy is worth flagging versus ignoring?

Maya, mid-level lending auditor
Maya
Age
34
Occupation
Mid-level Lending Auditor
Location
Toronto, ON

"The tool can show me exactly what it did. What nobody can tell me is whether I'm allowed to rely on it."

Goals
  • Complete lending audits faster without sacrificing defensibility
  • Understand what the agent is doing and why
  • Build a workflow she can rely on case after case
Pain points
  • No one defined which audit steps she can delegate to the agent
  • Output needs to be defensible to managers and regulators
  • Can't tell where the agent's judgment ends and hers begins
Current solution

Re-verifies everything the agent produces manually - erasing the time savings

Interest
  • Professional accuracy
  • Regulatory compliance
  • Efficient workflows
Solution

Calibration, not onboarding

A trust calibration layer that sits before the agent takes any action on a real case. Not onboarding. Onboarding teaches features. Calibration builds a working relationship. The core mechanic: the agent works a past resolved case the user already knows the answer to. She watches it reason through something familiar before relying on it for anything real.

Role and context intake
The agent asks before it acts. No configuration — just four plain-language questions.
Past case upload
She uploads a resolved case. The agent confirms its understanding. She corrects anything wrong.
Agent works the past case
She watches the agent reason through something she already knows. This is the baseline trust moment.
User evaluates and adjusts
She flags what it got right and what it missed. Preferences emerge from experience, not hypothesis.
Live case
The agent works the real case inside the context she's established. At judgment boundaries, it pauses.
User action
System response
Decision
What Changes

What AI actually mean for each team?

For Maya

She understands what the agent does before she relies on it. Her output stays defensible because she's never guessing what the agent did or why.

For IT

Per-team configuration is no longer an engineering dependency. The user builds her own workflow context in plain language, grounded in how she actually works.

For Leadership

AI usage measured by decisions made, not tokens consumed. Faster time to actual utility. Workflows that reflect real operations.

Stakeholders

What are their concerns?

StakeholderRolePrimary concern
Maya (end user)Lending auditorCan I trust this enough to use it on real cases?
Maya's managerAudit team leadIs the team's output still defensible and reviewable?
IT / implementationEnterprise deploymentCan we deploy without configuring every team individually?
Compliance / legalRisk oversightDoes agent-assisted output meet regulatory standards?
Procurement / leadershipBudget holderIs this producing value beyond token counts?
Prototype

See it in action

A walkthrough of the trust calibration onboarding flow - from intake questions to a saved agent context. Navigate through the screens using the buttons inside the prototype.

Open in full screen →

Limitations

What this doesn't solve

This calibration layer addresses the trust gap at the product level. It does not fix the organizational problem upstream - most companies buy AI tools without defining what good usage looks like for each team. Organizational guidance, training, and success metrics still matter and sit outside the scope of this design.

Success Metrics

How we'd measure this

MetricSignal
Completion rate of first-use flowAre users finishing calibration or dropping off?
Time to first real caseHow quickly does calibration convert to productive use?
User-reported confidence after step 3Does watching the agent work a past case build trust?
Review cadence on live casesAre users over-reviewing (low trust) or under-delegating?
Retention at 30 and 90 daysDoes early calibration produce sustained usage?
Open Questions

How to optimize this product?

  1. 1

    What happens when a user has no past case to upload? The flow needs a graceful fallback - a templated example or skip-ahead state - that doesn't break the trust-building sequence.

  2. 2

    How does the calibration layer interact with organizational compliance requirements? In regulated industries, the agent's reasoning may need to be logged regardless of user preference.

  3. 3

    At what point does a saved workflow need to be recertified? If regulations change or the user's role expands, stale preferences become a liability.

Next Steps

Where this goes next

Role-adaptive flow

The current prototype presents the same calibration experience regardless of intake answers. The next iteration uses those answers to adjust the flow: document types surfaced, judgment scenarios presented, and agent language throughout. An auditor and a sales ops analyst are doing fundamentally different work.

Multi-team rollout model

The current design solves for individual onboarding. The next question is organizational: how does a team lead deploy this across twelve auditors without each person starting from scratch? A shared workflow template, seeded by a team lead and adjustable per user, would reduce setup time and create consistency.

Workflow versioning

As regulations change or a user's role expands, saved preferences become stale. A lightweight versioning system that flags when a saved workflow hasn't been reviewed in a defined period would keep the calibration layer accurate over time.