AI Product Design
If you design AI features the same way you design ordinary software features, the result will likely disappoint users.
-
Ordinary software is deterministic—you click save, it saves, every time the same.
-
AI products are probabilistic—you ask it to write copy, it does great this time, but might fail next time.
When ordinary software fails, it shows an error message; when AI products fail, they might confidently make things up, making you think they're correct.
This fundamental difference means AI product design requires an entirely new approach.
Traditional product design pursues zero defects, where users get the same result every time. AI product design must acceptprobabilistic output, helping users understand and cope with the uncertainty of results.
AI Product Thinking
To design good AI products, you first need to understand the uniqueness of AI and the psychological changes it brings to users.
Probabilistic Output
The output of large models is essentially probability sampling. The same question may yield different answers each time — sometimes good, sometimes bad, sometimes mediocre. This is not a bug; it's a characteristic of AI.
For designers, this means several key principles:
| Design principles | Specific practices | Why it matters |
|---|---|---|
| Provide a retry mechanism | Let users "generate again" | Give users a second chance when they're unsatisfied |
| Show multiple options | Generate 3-5 versions at once for selection | Increase the probability of users finding a satisfactory result |
| Allow editing and modification | AI-generated content should be easy to edit | Users can correct imperfections |
| Set quality expectations | Tell users "results may need adjustment" | Avoid overly high user expectations |
Good AI products don't try to hide this uncertainty, but turn it into an advantage.
User Psychology: Expectation Management
User expectations of AI often swing between two extremes: either expecting too much, or completely distrusting it.
People using ChatGPT for the first time often marvel at how amazing it is, believing AI is omnipotent.
But when AI makes a silly mistake, they may immediately swing to "AI is just so-so."
The task of a product designer is toGuide user expectations into a reasonable range。
How?
-
First, clearly tell users what AI can and cannot do.
-
Second, when AI makes mistakes, don't try to hide them — face them honestly.
-
Third, give users control — let them adjust, edit, and override AI's output.
Good AI products make users feel: AI is my assistant, not my boss.
Failure Modes of AI Products
The way AI products fail is very different from traditional software.
Traditional software either works or doesn't — it crashes, reports errors, or features stop working.
AI product failures are more subtle:
| Failure mode | Symptoms | Response strategies |
|---|---|---|
| Hallucination | Fabricating facts, citations, and data that don't exist | Add fact-checking, provide source attribution, allow users to verify |
| Alignment failure | The answer doesn't match the user's intent | Provide clarifying questions, let users confirm intent, optimize through multi-turn dialogue |
| Unstable output quality | Sometimes very good, sometimes very bad | Provide multiple options, allow retries, let users rate and give feedback |
| Overconfidence | Saying something wrong in a confident tone | Add confidence display, use more cautious wording, encourage questioning |
| Context loss | Forgetting key information from previous conversations | Show context summaries, allow citing history, provide conversation memory management |
Understanding these failure modes is the first step in designing good AI products.
AI Feature Design Principles
Three core principles: progressive disclosure, user control, and transparent notification.
Progressive Disclosure
Don't pile all features in front of the user.
The ideal flow is: first give users a simple starting point, then gradually show more options as needed.
For example, when a user writes an email:
-
Step 1: Enter a subject or keywords, and the AI generates a draft.
-
Step 2: After reviewing the draft, the user can adjust tone, length, and style.
-
Step 3: If needed, they can further revise specific paragraphs or have the AI provide several different versions.
The benefits of this design are:
Beginners won't be intimidated by complex options, and experts can find enough control.
| Stage | What users see | User actions |
|---|---|---|
| Initial interface | Simple input box | Enter basic needs |
| After generation | Result + basic adjustment options | Choose "more formal", "more concise", etc. |
| Expand advanced options | Detailed parameter control | Adjust temperature, role, format, etc. |
Controllability: Letting Users Adjust AI Output
Users need to feel that they are the final decision-makers.
Provide control knobs, not a one-shot "take it or leave it" result.
Common control dimensions:
| Control dimension | Typical options | Applicable scenarios |
|---|---|---|
| Style/Tone | Formal, casual, humorous, professional | Writing, emails, copy |
| Length | Short, medium, detailed, long | Summaries, articles, reports |
| Complexity | Easy to understand, medium, professional depth | Explanations, tutorials, technical documentation |
| Creativity | Conservative, balanced, boldly innovative | Brainstorming, creative writing |
| Format | List, paragraph, table, outline | Notes, planning, documents |
More advanced control: let users directly edit the AI's prompt, or save their commonly used prompt templates.
Transparency: Informing Users This Is AI-Generated
Letting users know they are interacting with AI is not only an ethical issue, but also a product experience issue.
If users think it was written by a real person, they will feel deceived when they discover it is AI.
If it is stated from the beginning that it is AI-generated, users will actually be more forgiving and more willing to participate in improvements.
Several approaches to transparency:
First, clear labeling—use visual elements to distinguish AI content from user content.
-
Second, show the process—let users see how AI generates results (for example, showing the thinking process and retrieved sources).
-
Third, explain limitations—tell users what types of errors AI may make and how to identify them.
Transparency builds trust. When users know it is AI, they will use it in the right way—as reference rather than blind trust.
User Experience Design
AI products have several key experience points: loading states, error handling, and feedback mechanisms.
Loading State Design: Streaming Output
Large models take time to generate content. Making users wait idly for 10 seconds results in a poor experience.
The best practice isstreaming output—content appears on the screen word by word.
Why is it good?
-
First, users feel the system is working and hasn't frozen.
-
Second, users can start reading early and know halfway through whether it is the direction they want.
-
Third, if the direction is wrong, they can interrupt at any time without waiting for the full generation to finish.
Besides streaming output, you can also:
-
Show progress indicators—"thinking", "retrieving", "generating".
-
Provide estimated time—"approximately 10 seconds".
-
Provide a cancel button—allowing users to stop at any time.
Error Handling and Degradation Strategies
AI will inevitably make mistakes. Good product design makes errors less daunting.
When AI output is clearly problematic:
-
First, give users a simple "unsatisfied" or "retry" button.
-
Second, provide alternatives—"how about trying this angle?"
-
Third, allow users to easily roll back to a previous state.
More severe errors (e.g., the model completely fails):
Have a fallback plan. For example, when AI is unavailable, provide a template library for users to choose from manually.
| Error severity | User behavior | Product response |
|---|---|---|
| Minor issue | Result is slightly flawed but usable | Provide editing features for users to fine-tune |
| Moderate issue | Result is incorrect, needs regeneration | Provide a retry button, or suggest adjusting the input. |
| Critical issue | AI not working at all | Provide a fallback solution, such as a template library |
Feedback Mechanisms
User feedback is a valuable resource for improving AI products.
But the feedback feature shouldn't be too complex, otherwise users won't want to use it.
Simple and effective feedback design:
-
First, like/dislike — one-click expression of satisfaction or dissatisfaction.
-
Second, a brief multiple-choice — "What's wrong?" (Too long, too short, off-topic, factual errors...).
-
Third, an optional text box — allows users to elaborate on the issue, but it's not required.
-
Fourth, tell users what the feedback is for — "Your feedback will help us improve."
After feedback is collected, it's best to give users a confirmation in the interface — "Thank you for your feedback, we've received it."
Prompt Productization
Good prompts are the core competitive advantage of AI products. Turn prompts from "magic" into manageable product features.
Packaging Prompts as Product Features
Ordinary users don't need to know what a "system prompt" is.
They only need to know: click this button, and you get a draft of a formal email.
Therefore, the designer's job is to encapsulate complex prompts into simple feature buttons.
For example:
The original prompt may be very long:
你是一个专业的邮件写作助手。请帮用户写一封正式、礼貌、简洁的商务邮件。 要求:1. 开头要有合适的称呼;2. 正文表达清晰;3. 结尾要有礼貌的结束语。 语气要专业但不生硬,友好但不过分随意。
After productization, what users see is just:
-
A button — "Write a business email."
-
A few simple options — "Formal/Neutral/Friendly", "Brief/Detailed".
This is the core of prompt productization:Keep complexity to yourself, keep simplicity for users.。
Version Management of System Prompts
A prompt isn't done once it's written; it needs continuous iteration.
You may find that a new version of the prompt works better for scenario A, but worse for scenario B.
Therefore, prompts need version management like code.
Key practices:
-
First, give each version a number or name — "v1.0", "v1.1", "Experimental - Friendlier".
-
Second, record changes for each version — "What changed, why it changed, expected impact."
-
Third, you can run multiple versions simultaneously for A/B testing.
-
Fourth, you can quickly roll back to a previous version.
When iterating on prompts, always retain the ability to roll back. A new version may bring unexpected problems.
A/B Testing Prompts
Two versions of a prompt, which is better? Don't guess; let the data speak.
The idea of A/B testing:
-
Some users use version A, and some users use version B.
-
See which version has higher user satisfaction, lower retry rate, and better completion rate.
Below is a simple Python script demonstrating how to perform A/B test analysis for prompts:
Examples
# Prompt A/B test analysis script
# Used to compare the effectiveness of two prompt versions
# ============================================
from dataclasses import dataclass
from typing import List, Dict, Optional
import statistics
@dataclass
class TestResult:
"""Data structure for a single test result"""
prompt_version: str # "A" or "B"
user_satisfaction: int # User satisfaction 1-5
retry_count: int # Number of user retries
task_completed: bool # Whether the task is completed
time_spent_seconds: int # Time spent
test_case_id: str # Test case identifier (e.g., different query types)
def analyze_ab_test(results: List[TestResult]) -> Dict:
"""Analyze A/B test results and return comparative statistics"""
# Group by version
group_a = [r for r in results if r.prompt_version == "A"]
group_b = [r for r in results if r.prompt_version == "B"]
if not group_a or not group_b:
return {"error": "Need test data from at least two versions"}
def calc_stats(group: List[TestResult]) -> Dict:
"""Calculate statistics for a single group"""
satisfaction_scores = [r.user_satisfaction for r in group]
return {
"sample_size": len(group),
"avg_satisfaction": statistics.mean(satisfaction_scores),
"median_satisfaction": statistics.median(satisfaction_scores),
"avg_retry": statistics.mean([r.retry_count for r in group]),
"completion_rate": sum(1 for r in group if r.task_completed) / len(group),
"avg_time": statistics.mean([r.time_spent_seconds for r in group]),
}
stats_a = calc_stats(group_a)
stats_b = calc_stats(group_b)
# Compare key metrics
comparison = {
"satisfaction_diff": stats_b["avg_satisfaction"] - stats_a["avg_satisfaction"],
"retry_diff": stats_b["avg_retry"] - stats_a["avg_retry"],
"completion_diff": stats_b["completion_rate"] - stats_a["completion_rate"],
}
# Can also break down by test case (e.g., different types of queries)
by_test_case = {}
test_case_ids = set(r.test_case_id for r in results)
for case_id in test_case_ids:
case_results = [r for r in results if r.test_case_id == case_id]
case_a = [r for r in case_results if r.prompt_version == "A"]
case_b = [r for r in case_results if r.prompt_version == "B"]
if case_a and case_b:
by_test_case[case_id] = {
"a_score": statistics.mean(r.user_satisfaction for r in case_a),
"b_score": statistics.mean(r.user_satisfaction for r in case_b),
}
return {
"version_a": stats_a,
"version_b": stats_b,
"comparison": comparison,
"by_test_case": by_test_case,
"winner": _determine_winner(stats_a, stats_b),
}
def _determine_winner(stats_a: Dict, stats_b: Dict) -> Optional[str]:
"""Determine which version is better based on statistics"""
# Consider satisfaction, completion rate, and retry count comprehensively
a_score = (
stats_a["avg_satisfaction"] * 0.5 +
stats_a["completion_rate"] * 10 * 0.3 +
(5 - stats_a["avg_retry"]) * 0.2
)
b_score = (
stats_b["avg_satisfaction"] * 0.5 +
stats_b["completion_rate"] * 10 * 0.3 +
(5 - stats_b["avg_retry"]) * 0.2
)
# If the difference is too small, it's inconclusive
if abs(a_score - b_score) < 0.1:
return "tie" # Tie
return "A" if a_score > b_score else "B"
def print_report(analysis: Dict) -> None:
"""Print a readable A/B test report"""
print("=" * 60)
print("EXAMPLE Prompt A/B Test Report")
print("=" * 60)
if "error" in analysis:
print(f"Error: {analysis['error']}")
return
print(f"\nVersion A (baseline):)
print(f" Sample size: {analysis['version_a']['sample_size']}")
print(f" Average satisfaction: {analysis['version_a']['avg_satisfaction']:.2f}/5")
print(f" Completion rate: {analysis['version_a']['completion_rate']:.1%}")
print(f" Average retries: {analysis['version_a']['avg_retry']:.2f}")
print(f"\nVersion B (new approach):)
print(f" Sample size: {analysis['version_b']['sample_size']}")
print(f" Average satisfaction: {analysis['version_b']['avg_satisfaction']:.2f}/5")
print(f" Completion rate: {analysis['version_b']['completion_rate']:.1%}")
print(f" Average retry count: {analysis['version_b']['avg_retry']:.2f}")
print(f"\n"Comparison results:")
diff = analysis["comparison"]
print(f" Satisfaction change: {diff['satisfaction_diff']:+.2f}")
print(f" Completion rate change: {diff['completion_diff']:+.1%}")
print(f" Retry count change: {diff['retry_diff']:+.2f}")
winner = analysis["winner"]
if winner == "tie":
print("\n"Conclusion: The two versions perform similarly; more data or adjustments are needed.")
else:
print(f"\n"Conclusion: Version {winner} performs better.")
if analysis.get("by_test_case"):
print(f"\n"Breakdown by test case:")
for case_id, scores in analysis["by_test_case"].items():
better = "A" if scores["a_score"] > scores["b_score"] else "B"
print(f" {case_id}: A={scores['a_score']:.2f}, B={scores['b_score']:.2f} ({better} better)")
# ============================================
# Usage example
# ============================================
if __name__ == "__main__":
# Simulate some test data
test_data = [
# Test results for Version A
TestResult("A", 4, 0, True, 15, "Email writing"),
TestResult("A", 3, 1, True, 25, "Email writing"),
TestResult("A", 5, 0, True, 12, "Summary generation"),
TestResult("A", 2, 2, True, 40, "Summary generation"),
TestResult("A", 4, 0, True, 18, "Code explanation"),
TestResult("A", 3, 1, False, 35, "Code explanation"),
# Test results for Version B
TestResult("B", 5, 0, True, 12, "Email writing"),
TestResult("B", 4, 0, True, 18, "Email writing"),
TestResult("B", 4, 0, True, 10, "Summary generation"),
TestResult("B", 3, 1, True, 25, "Summary generation"),
TestResult("B", 5, 0, True, 15, "Code explanation"),
TestResult("B", 4, 0, True, 20, "Code explanation"),
]
# Analyze and print report
analysis_result = analyze_ab_test(test_data)
print_report(analysis_result)
Run this script, and you'll get a clear A/B test report showing which version of the prompt performs better.
============================================================ EXAMPLE Prompt A/B 测试报告 ============================================================ 版本 A (基准): 样本量: 6 平均满意度: 3.50/5 完成率: 83.3% 平均重试次数: 0.67 版本 B (新方案): 样本量: 6 平均满意度: 4.17/5 完成率: 100.0% 平均重试次数: 0.17 对比结果: 满意度变化: +0.67 完成率变化: +16.7% 重试次数变化: -0.50 结论: 版本 B 表现更好 按测试用例细分: 邮件写作: A=3.50, B=4.50 (B更好) 摘要生成: A=3.50, B=3.50 (A更好) 代码解释: A=3.50, B=4.50 (B更好)
AI Feature Prototyping
Before investing significant development resources, first use a low-cost approach to verify whether the AI feature is actually useful.
Rapid Prototyping with Claude/ChatGPT
The simplest prototype: directly use the large model chat interface to simulate product functionality.
For example, you want to build an interview question generator:Write a prompt describing the feature you want, then input some test cases and observe the output quality.
The goal of this stage is to verify:
-
First, can the AI do this task?
-
Second, is the output quality good enough?
-
Third, would users find this feature useful?
No need to write code or design a UI; just use plain-text conversation to test the core value.
If this stage doesn't perform well, either improve the prompt or reconsider whether this feature is worth building.
AI Prototype Components in Figma
After confirming the core functionality is valuable, the next step is to create a high-fidelity prototype.
Figma is a great tool for this.
You can design:
-
How does the user input?
-
How is the AI output displayed?
-
How do users adjust and edit?
The key issimulating real interaction flows, not just drawing static interfaces.
You can use Figma's interaction features to "generate" different content when users click buttons (prepare several example results in advance).
This way users can feel: Oh, so this is how the product works.
Low-Code AI Application Tools
If you need a more realistic experience, you can use low-code tools to quickly build a working prototype.
These tools let you connect a large model to a real web application without writing backend code.
The benefits are:
Users can actually operate it, not just watch a demo.
You can collect real usage data and feedback.
After successful validation, you can use this prototype as a reference for development.
The goal of a prototype is to learn, not to be perfect. Validate ideas as quickly as possible, discover problems, and iterate.
Evaluating AI Output Quality
How do you judge whether AI output is good? You need an evaluation system that combines qualitative and quantitative methods.
Qualitative Evaluation: Human Review
Some things can only be judged by humans.
For example: Is the answer helpful? Is the tone appropriate? Is the logic coherent?
The approach for human review:
-
First, design a scoring dimension table—for example, relevance, accuracy, usefulness, and safety.
-
Second, have multiple reviewers score independently to reduce personal bias.
-
Third, organize the scores and comments into actionable improvement suggestions.
Human review is important, but it is slow, expensive, and difficult to scale. So it needs to be combined with automated evaluation.
Quantitative Evaluation: Automated Assessment
Use programs to evaluate AI output.
Common automated evaluation methods:
| Method | Approach | Applicable scenarios |
|---|---|---|
| Rule-based checking | Check keywords, format, length, and whether sensitive content is included | Basic quality control |
| LLM self-evaluation | Use another large model to evaluate output quality | Open-ended tasks |
| Semantic similarity | Compare the semantic closeness of the output to a reference answer | Tasks with standard answers |
| User behavior metrics | Track retry rate, completion rate, and satisfaction score | Ongoing monitoring after launch |
Automated evaluation cannot fully replace humans, but it can:
-
First, quickly filter out obviously problematic outputs.
-
Second, run regression tests during version iterations — "did this change make any scenarios worse?"
-
Third, monitor online service quality and identify issues promptly.
Gold Test Set Construction
Collect a set of high-quality test cases, and run them on this set every time you modify prompts or models.
This is thegolden test set。
A good test set should:
-
First, cover typical usage scenarios — common types of user requests.
-
Second, include edge cases — extreme, error-prone inputs.
-
Third, have reference answers or scoring criteria — know what counts as a "good" output.
-
Fourth, be of moderate size — tens to hundreds of cases, representative enough without running too slowly.
The golden test set is like a safety net: you can iterate with confidence without worrying about accidentally breaking important scenarios.
Commercialization Paths for AI Products
Once you have a good product, how do you generate revenue from it? There are several common commercialization models for AI products.
| Model | Approach | Typical examples | Advantages | Challenges |
|---|---|---|---|---|
| Subscription | Pay monthly/yearly | ChatGPT Plus | Stable, predictable revenue | Need to continuously deliver value |
| Usage-based billing | Pay based on usage (per call or token) | OpenAI API | Low barrier for users, on-demand usage | Users may worry about cost |
| Freemium | Basic features free, premium features paid | Notion AI | Easy to acquire users | Strike the right balance between free and paid |
| Embedding in enterprise products | Selling AI capabilities to enterprise customers | Customer service bots, document assistants | High average order value, long decision cycle | Requires sales and customer success teams |
Which model to choose depends on your product positioning and target users.
For general consumers, freemium or subscription models are more common.
For enterprise customers, usage-based pricing or customized solutions are more suitable.
But regardless of the model, the core principle is:Users must be able to genuinely feel the value。
Case Study: Deconstructing Excellent AI Products
Let's look at a few well-made AI products and learn from their design thinking.
ChatGPT: A Model of Conversational Interaction
ChatGPT's success is not only a technological success, but also a design success.
Its core design decisions:
-
First, the conversation format — users don't need to learn a new interaction method; it's like chatting with a person.
-
Second, streaming output — characters appear one by one, making the user feel like the AI is thinking.
-
Third, context memory — you can say "make that last part shorter," and it remembers what "that last part" was.
-
Fourth, simple retry — not satisfied? Just click "regenerate."
-
Fifth, conversation history management — you can go back to previous conversations and continue where you left off.
These designs look simple, but together they create a smooth experience.
GitHub Copilot: Embedding AI into the Workflow
Copilot doesn't try to "replace programmers," but rather "augment programmers."
Its design wisdom:
-
First, non-intrusive — the AI works in the background, proactively offering suggestions while you write code without disrupting your thinking.
-
Second, low commitment — you can view the AI's suggestion; if you like it, press Tab to accept it; if not, keep typing.
-
Third, context awareness — it looks at your code, comments, and file names to understand what you're trying to do.
-
Fourth, multiple options — it provides several suggestions, and you can pick the one that best matches your intent.
Copilot's success shows:The best AI product makes users not feel the AI's presence, only feel that they've become more powerful.。
Notion AI: AI as a Writing Copilot
Notion AI didn't create a standalone product; instead, it embedded AI capabilities into the existing note-taking tool.
Its design highlights:
-
First, context is your document — you don't need to copy and paste; the AI works right where you're writing.
-
Second, preset templates — "continue writing," "summarize," "make it more formal," common actions triggered with one click.
-
Third, mixed editing — you can modify the AI-generated content at any time, and the AI can continue writing based on your edits.
This design makes the AI part of the creative process, rather than a standalone tool.
Other extensionsFrom these three cases, we can see: good AI products are not about showing off technology, but about truly solving problems and making users' work easier.