logo/All Projects/Amazon Bedrock
AWS·Generative AI Platform UX·2024

Amazon Bedrock

Designing scalable workflows and reducing ambiguity for GenAI applications.

Role
UX Designer (AWS GenAI)
Scope
Bedrock Flows · Model Evaluations
Partners
PM · Engineering · External feature partners · Design system team
Status
Publicly launched
Amazon Bedrock FlowsAmazon Bedrock Flows execution detailsAmazon Bedrock workflow detailBedrock Model Evaluations
Amazon Bedrock Flows and Model Evaluations — execution-trace visibility, input validation, and result clarity.

Overview

Amazon Bedrock is AWS's managed platform for building, evaluating, and operating generative AI applications using foundation models. I led UX for two core platform capabilities: Bedrock Flows (async executions and trace visibility) and Bedrock Model Evaluations (errors, input validation, and result clarity). Both shipped publicly.

The problem

Customers could design and test Flows synchronously, but real production use cases needed long-running executions, the ability to re-run trusted flows without re-authoring, and clear visibility into progress, failures, and intermediate steps. Model evaluations had a parallel problem: cryptic errors at start, malformed inputs that wasted compute, and evaluation results that were hard to interpret — especially for RAG pipelines.

"How do we scale GenAI workflows beyond synchronous limits — and improve trust and observability when they run?"

Approach: Bedrock Flows

The work followed a discovery-first arc: PM alignment, conversations with feature partners and customers, then design through internal review to public launch.

  • Async executions. Identified asynchronous executions as a critical platform gap and designed UX for creating executions from previously validated flows.
  • Execution monitoring. Designed UX for monitoring execution status over time and inspecting trace details for each step in a flow execution.
  • Launch. Opened designs for broad internal review, incorporated feedback, finalized high-fidelity designs, and supported public launch.

Approach: Model Evaluations

Model evaluations require structured datasets — JSONL files — and are typically used for comparing models or evaluating RAG pipelines. The brief was to make the experience legible at every failure point.

  • Error states and guidance. Designed improved errors and inline guidance during evaluation setup to reduce confusion and failure rates.
  • Input validation. Created UX for validating uploaded JSONL files before evaluations run — preventing wasted compute and time.
  • Results clarity. Designed clearer evaluation result views with scannable presentation and source-citation support for RAG evaluations.

Outcome

Both features shipped as publicly available Bedrock capabilities. Async executions enabled production-scale GenAI workflows beyond synchronous limits. Execution tracing improved observability and trust in AI workflows. Better error handling and input validation reduced friction during evaluation setup, and the results UX made model evaluations more actionable and trustworthy.

Reflections

  • Trust through transparency. In AI systems, users need visibility into what's happening. Execution traces and clear status indicators build confidence.
  • Proactive error prevention. Validating inputs before expensive operations run saves time and reduces frustration. Guide users toward success.
  • Partner-driven discovery. Direct conversations with feature partners and customers surface real-world needs that internal assumptions miss.

This case study is described at a high level. Some details, metrics, and internal tooling are generalized or omitted under confidentiality. A deeper walkthrough with workflows, error states, and execution traces is available on request.

Want the deeper walkthrough?

I'm happy to share workflows, error states, and the execution-trace UX in detail.

Get in touch →