September 1, 2026



September 1, 2026



AI has enabled research to scale at a level we’ve never seen before. But is it possible to maintain a high level of quality at that scale? What gets lost in the process, and how can researchers and designers continue to be accountable for the work that’s produced?
In the final installment of this DXC and Dscout webinar series, our panelists discuss where AI can help, what its limitations are, and what we need to keep an eye on as humans at the helm.
Panelist participants included:
You can watch the webinar in its entirety here.
Traditional user experience evaluation relied heavily on deterministic software interfaces.
Researchers could conduct heuristic evaluations, run usability tests with a handful of participants until reaching saturation, or rely on QA teams to verify that specific UI buttons and flows yielded identical results every time.
In deterministic software, input directly equaled output, allowing teams to easily verify whether a bug fix or new release was objectively better than the last.
Generative AI fundamentally disrupts this paradigm because LLMs are non-deterministic and stochastic systems. When two users enter the same prompt or navigate the same AI feature, they will likely receive distinct outputs, meaning classical testing methods no longer reach saturation or capture every potential outcome.
As a result, product and research teams need to shift toward continuous evaluation models that assess prompt chains, guardrails, and conversation flows against fundamental quality requirements to determine if a system is truly ready to ship.
Many large organizations operate without a unified, shared definition of what high quality looks like in design and research. Instead, quality decisions are often dictated by individual egos, subjective opinions, departmental debates, and unwritten institutional knowledge.
This fragmentation leads to constant friction—such as endless arguments over design system usage or brand tone compliance—making consistent execution across massive enterprises nearly impossible.
AI offers a powerful solution to this ambiguity by functioning as an objective gatekeeper when architected correctly within an organizational system, harness, or factory. By encoding precise accessibility rules, brand guidelines, and governance policies directly into system architecture, teams can catch design drift before it reaches production.
Rather than trusting AI to make autonomous decisions, setting rigorous baseline standards allows AI to handle initial policy enforcement, freeing human designers to focus their energy on higher-level strategic problems.
"If we set up our harness properly...and if we have the right tools, the right architecture, the right connection points, the right agents, the right guardrails, we can catch the drift before it starts." — Christina Vallery, Chief Design Officer at The Cigna Group
Historically, organizational understanding of "good" design lived abstractly within vague design principles, scattered catalogs of past work, or simply inside the heads of senior practitioners.
While human teams could rely on implicit intuition to navigate these uncodified gaps, generative AI systems possess no natural instinct for company-specific context. To generate meaningful outputs, AI must be explicitly taught what constitutes quality and success within a particular organization's customer domain.
It is important to remember that any agentic research solution can only be as good as the knowledge it's built on.
Building effective agentic research systems requires moving past tactics like standalone synthetic personas, which are ultimately meaningless without rich context.
Instead, organizations must connect disparate customer insight channels—including research repositories, customer support tickets, and direct feedback signals—into a centralized data foundation.
Unifying these disconnected inputs transforms implicit institutional knowledge into structured, machine-readable institutional insight that elevates both automated workflows and human decision-making.
“We are focusing on connecting different sources of customer insight across our own research repository, customer tickets, etc. As we connect them, we hope to build a much more living understanding of what's going on with our customers. Over time, we want to expand that by adding more signals from across the business.” — Anshuman Kumar, SVP, Head of Design at Datadog
Evaluating digital user experience systematically requires breaking down customer touchpoints into four distinct, interconnected layers:
To maintain these standards across AI-generated experiences, organizations need to translate their human design guidelines into machine-readable assets. By creating structured markdown files, explicit agent rules, and centralized prompt libraries, teams enable AI agents to automatically reference approved standards during generation.
This structural foundation prevents individual designers and developers from having to recreate guardrails from scratch, embedding design consistency directly into the production lifecycle.
Training generative AI models resembles machine teaching, where initial model training functions like a lecture, and evaluation serves as the exam. To evaluate whether an AI agent meets production standards, teams should develop highly detailed, explicit rubrics that define exactly what success looks like in specific operational scenarios.
For example, an AI moderation tool designed to conduct unmoderated research follow-ups requires precise scoring criteria to balance active listening without interrupting users or lagging excessively.
Once established, these rubrics enable teams to engage in a technical optimization process known as hill climbing. By running automated scenarios against structured rubrics, product teams can iteratively refine their prompt strategies and model parameters to systematically increase pass rates from baseline levels up to production-ready thresholds.
“Now the question becomes: how do you create a good rubric, and who should create it? That's a massive opportunity for UX teams. Often, they come from engineering teams as very technical rubrics. They might define what is technically a good output, but that might not be at all what a user thinks is a good output.” — Michael Winnick, CEO at Dscout
High-quality content design extends far beyond applying basic voice and tone guidelines; it requires balancing business objectives, back-end system constraints, regulatory policies, and evolving customer expectations.
Traditionally, navigating these complex, multifaceted trade-offs relied heavily on the implicit mental models of experienced content strategists. Because these nuances lived almost entirely inside individual practitioners' heads, specialized work remained constrained to specific experts with deep institutional knowledge.
The push toward AI automation is forcing organizations to systematically expose and document this implicit knowledge. Building effective automated workflows requires defining semantic layers, mapping existing customer journeys, and encoding policy rules so machines can process the same context that human experts use.
This transition elevates the visibility of content design, transforming previously insular, hidden expertise into scalable assets that benefit the entire organization.
"It's actually a really thrilling time to be able to expose what has traditionally been a very insular kind of experience." — Catherine Walker, Head of Content Strategy and Architecture at The Cigna Group
The phrase "human in the loop" inadvertently positions human practitioners as passive bystanders sitting alongside automated systems, rather than active owners. Reframing this dynamic to "human at the helm" re-establishes humans as the primary creators, commanders, and authors of AI workflows.
This language shift emphasizes that accountability and strategic control remain entirely with human leaders, preventing organizations from abdicating operational responsibility to automated software.
Operationalizing a "human at the helm" philosophy involves establishing a clear three-tiered governance framework:
"The very idea of a human just being in the loop with AI assumes that the human is actually a bystander and not an owner, and I really want to change the vocabulary there. The term that we've been using internally is human at the helm."
— Christina Vallery, Chief Design Officer at The Cigna Group
As generative AI makes it drastically easier to build and ship digital experiences, organizations need to balance speed with rigorous quality enforcement. Without clear experience benchmarks, teams risk releasing poorly designed tools simply because they’re technically functional.
Establishing formal evaluation scorecards ensures that products undergo objective quality assessments before reaching consumers, holding teams directly accountable for the end-user experience.
One example of this accountability: an internal chatbot, developed outside standard design channels, was evaluated against an experience scorecard using structured rubrics. When the tool produced a failing score across simplicity and content metrics, the organization halted its public release.
Enforcing quantitative quality bars prevents substandard products from going to market—and sparks necessary conversations around experience standards early in product development.
Modern AI architectures frequently employ online evaluations, where "LLM-as-a-judge" evaluators monitor system performance in real time. These automated monitors continually assess outputs against basic rules—such as checking whether a survey question generator accidentally creates double-barreled questions—and generate scores to trigger corrective actions, or track overall system health without requiring manual intervention.
However, automated evaluators cannot resolve every complex edge case or subjective judgment call. Technical "human in the loop" mechanisms are explicitly designed to identify when system confidence drops, or when an interaction exceeds automated parameters, triggering a flag for human review.
Establishing clear boundaries for semi-autonomous operations ensures that human expertise is deployed precisely where complex reasoning is necessary.
"To a degree, all of these systems face these questions. If you talk about what is agentic, we could have a whole debate around that. But these questions around when is the system performing semi-autonomously? That's what you're trying to fine-grain in some of these decisions about when you pull a human in or not."
— Michael Winnick, CEO at Dscout
Traditional user research historically relied on data compression, flattening rich human variability into generalized user personas, market segments, and statistical averages.
While necessary due to time and resource constraints, this compression often stripped away individual nuance, opening up significant equity and accessibility gaps. By designing for an abstract average user, organizations frequently overlooked edge cases and specialized needs that didn’t fit neatly into broad buckets.
Agentic research workflows unlock a "post-persona" paradigm, enabling teams to evaluate human variability at unprecedented scale and depth. Instead of relying on static persona proxies, AI agents allow researchers to test experiences against thousands of unique accessibility profiles, situational constraints, and personal contexts simultaneously.
This capability allows designers to map the full topology of human need and create truly personalizable experiences that respond to individuals in real time.
Despite the immense analytical power of agentic tools, AI systems possess fundamental limitations when interpreting human behavior. While agents can process massive repositories of survey data, behavioral telemetry, and interview transcripts, they lack lived human experience, empathy, and theory of mind. Consequently, an AI agent can’t intuitively grasp the underlying emotional resonance or social context that gives user feedback its deeper meaning.
Human researchers remain essential because they bring "qualitative validity"—the distinct ability to discern whether an observed user issue represents an isolated occurrence or a fundamental, generalizable principle.
While automated systems excel at surface-level pattern recognition, determining whether a finding warrants strategic action or requires broader quantitative validation still demands human craft, strategic judgment, and empathetic reasoning.
"What an agent lacks is a lived experience and the inherent ability to apply empathy and theory of mind to a data set that would allow it to fully grasp an insight, and know that it has more generalizable power." — Christian Rohrer, VP of Experience Design at TD Bank Group
Scaling quality with AI products doesn’t revolve around a single solution, but multi-pronged efforts headed by institutional leaders.
Not only does it require a strong point of view and philosophical underpinning, but it’s also necessary for everyone to adapt to the ground shifting underneath them and create explicit definitions of what defines quality.
This process is forcing institutional knowledge to the surface, and requiring new ways to assess standards within increasingly complex and customized systems. At the end of the day, human judgment is still one of leaders’ most invaluable and indispensable tools.