Case study · Insight · AI product process
Living Spec: How Our Product Team Learned to Build AI Features
- Overview
- When our organization started building AI features, no team had done it before. Nobody knew what the model needed, who decided how it should behave, or what should happen when it was wrong. I proposed one living spec that holds everything. Each team claims its part, AI helps draft each one, and every decision is written back until launch.
- Role
- Product Designer · author of the workflowI don't write code. I designed the process from the design and product side: the spec template, the questions every AI feature has to answer, and the review gates. Piloted on Nia, then proposed to the product org.
- Keywords
- AI product processCross-team collaborationModel Context ProtocolAI governance
- Status
- Piloted on Nia · proposed to the product org · still evolving
spec per AI feature, where every team claims its part
of the Figma first pass drafted by AI through MCP, from design system components
phases, each closed by a named human sign-off
change in time from brief to first design review
01 · The problem
Nobody knew how to build an AI feature.
Nia was our first AI feature, and the brief was "let's add AI to the platform." Product, design, engineering, data, and QA each owned a piece. None of us had built something whose answers change, depend on data, and can be wrong. Every team had questions that no one owned.
How do we measure success for something that talks?
How should it behave, and what should it never do?
Which model, and how does it plug into checkout?
Where do the answers come from, and are they current?
How do you test an answer that's different every time?
What is it allowed to say about products and prices?
✕ no shared plan · the answers lived in vendor PDFs, meetings, Slack, and people's heads
No playbook
AI features add questions regular features never had: behavior, data, failure, and evaluation. We had templates for none of them.
Everyone waited on everyone
Each team assumed another owned the hard questions, so they surfaced late, when they were expensive to answer.
Knowledge was scattered
Every new teammate, and every AI prompt, started from zero because the context wasn't written down in one place.
02 · The insight
The missing piece wasn't code. It was one plan every team could read.
I'm a UX designer with a product mindset, and I don't write code. What I can do is turn a messy problem into clear questions and decisions. That turns out to be exactly what AI needs too: plain language, with context. So I designed the workflow around a document, not a tool.
If every question has a section, and every section has an owner, nothing falls through.
One spec holds everything
Evidence, behavior, data, guardrails, decisions. If it isn't in the spec, it doesn't exist.
Every section has an owner
Teams claim their part at kickoff. No question stays unowned.
Plain language first
Anyone can read it, and so can AI. Owning a decision never requires code.
03 · One spec, every team
Every team claims its part of the same file.
Each AI feature starts from the same template. Sections tagged "new for AI" are the questions regular features never had. Pick a team to see what it owns.
Problem & success
What are we solving, and how will we know it worked?
User evidence
What do users actually say and do? Every claim links to a source.
AI behavior
What does it do, how does it sound, and what will it never do?
Data & sources
Where do answers come from, and how fresh are they?
Architecture
Which model, and how does it connect to cart, checkout, and accounts?
Failure & fallback
What happens when it's unsure, wrong, or unavailable?
Guardrails
What is it not allowed to say or do, and who approves the limits?
UI & components
Which design system parts does it use? AI drafts only from these.
Evaluation
Which test prompts must it pass, and what counts as a good answer?
Decision log
What we decided, why, who signed it, and the evidence behind it. Every team writes here.
10 sections, 6 of them new for AI features. Every section has exactly one owner.Product owns 1 section and contributes to 4.Design owns 4 sections and contributes to 4.Engineering owns 1 section and contributes to 5.Data owns 1 section and contributes to 2.QA owns 1 section and contributes to 2.Security & legal owns 1 section and contributes to 1.
04 · How work moves
Five phases. AI drafts, people sign.
The spec starts from two inputs and grows through five phases. Select a phase to see what AI drafts, what a person owns, what the spec gains, and who signs off.
partner docs, platform requirements
research, NPS, tickets, analytics
One file per feature. The single source of truth from research to launch.
phases · each reads from and writes back to the spec
Clusters NPS comments, support tickets, and analytics into themes, each linked to its source.
Product and design validate the themes, pick the problem worth solving, and define success.
§1 Problem & success§2 User evidence
Problem statement agreed before any design starts.
✓ signed by ProductFirst-pass screens and conversation flows in Figma through MCP, built only from design system components. Lists every state, including when the AI is unsure.
How the AI behaves, the flow, the hierarchy, and what it must never do. Deciding what not to build.
§3 AI behavior§6 Fallback§8 Components
Design review with product, engineering, and data.
✓ signed by DesignCoding agents read the same plain-language sections. Drafts tickets straight from the spec.
Engineering owns architecture and model integration, data owns sources. Design answers open questions in the spec, not in Slack.
§4 Data & sources§5 Architecture
Scope confirmed against the spec before the sprint starts.
✓ signed by Eng leadGenerates test prompts and expected answers from the evaluation section, and drafts usability test scripts.
QA judges the answers, not only the click paths. Design runs usability sessions and decides what changes.
§9 Evaluation§10 Decision log
Every test prompt passes, or has a logged exception.
✓ signed by QADrafts release notes and help content, then summarizes post-launch metrics.
Product makes the go / no-go call and security confirms the guardrails. Post-launch data flows back in as new evidence.
§7 Guardrails§2 User evidence
Go / no-go review.
✓ signed by Product + Security↺ every decision is written back into feature-spec.md, until launch
05 · Inside a real spec
Plain language, with an owner on every section.
This is roughly what Nia's spec looked like mid-design. No code anywhere. Engineers, QA, and AI agents all read the same text.
---feature: nia-persistent-sidebarstatus: design · v0.7sources: vendor-input.md, nps (500+ verbatims), usability-r1--- ## Problem @product1Buyers leave the product page to ask Nia a question,then lose their place. ## AI behavior @design2- Suggests products from the buyer's profile- Adds to cart only after the buyer confirms- Never guesses compatibility it can't source ## Data & sources @data- Specs, pricing [src: product catalog]- Buyer context [src: account profile] ## Fallback @design @eng3- Not sure → say so, link to the product page- Checkout issue → hand off to manual checkout ## Decisions @everyone4- v0.6 Modal → persistent sidebar why: keep browsing context while chatting owner: @design · src: usability-r1 · ✓ product ## Evaluation @qa5- "Does this laptop work with our dock?" → cites the spec sheet, or says it isn't sure- Chat history persists across page navigation
An owner on every section
The tag says who answers it. At kickoff, an empty section with no owner is the first red flag.
Behavior, written by design
How the AI acts is a UX decision. Writing it in plain sentences made it something product, engineering, and security could all review.
Failure is designed up front
Every AI feature is wrong sometimes. The fallback section makes "what happens then" a decision, not a surprise.
Decision log
What, why, who, and the evidence. Highlighted: the modal-to-sidebar pivot from usability testing.
Evaluation in questions
QA tests answers, not just click paths. Each test prompt comes with what a good answer looks like.
06 · Designed without code
I write in plain language. AI does the translating.
Every part of the workflow a designer touches is words and structure. AI turns those words into drafts, and a person checks each one. My leverage is the template, the questions it forces, and the judgment at each gate.
AI drafts about 80%. The other 20% is the design job.
AI drafts
- First-pass layouts from design system components
- Every state: empty, loading, error, unsure
- Copy variants and edge-case lists
- Research themes clustered from NPS and tickets
- Tickets and test cases from acceptance criteria
The designer decides
- Which problem is worth solving
- How the AI behaves, and where it stops
- Trade-offs between user needs and business goals
- What not to build, and what moves to long term
- Accessibility, brand fit, and final sign-off
07 · Enterprise guardrails
Speed only counts if legal, security, and brand can trust it.
A process that works for one designer breaks in a Fortune 500 company. These guardrails are what made it something the org could say yes to.
Approved models only
AI runs on enterprise-approved, in-tenant endpoints. Nothing from the spec goes into consumer tools.
Sanitize before merge
Customer PII and partner-confidential terms are stripped from inputs before they enter the spec.
The design system is the fence
Through MCP, AI composes only from published components and tokens. A new component needs design review.
A named owner at every gate
AI never approves anything. Each phase closes with a person signing off inside the spec.
Traceable by default
Every decision records what changed, why, who decided, and the evidence behind it.
Versioned
Every change to the spec has an author and a history, and can be rolled back.
08 · The pilot
We ran it on Nia first.
Nia, our GenAI shopping agent, was the first feature built this way. It was a good test: new territory, a vague brief, and five teams that had never built AI together.
500+ NPS comments, clustered by AI
We fed NPS results from Power BI into Azure OpenAI to find the most common pain points, then validated them with product and engineering in a Lean canvas workshop.
Every team claimed its section
Data and engineering sourced and structured product specs, pricing, and user data so Nia's answers had a real foundation. Design owned behavior and fallback.
First-pass screens drafted through MCP
AI drafted the Figma first pass from our design system. I spent my time on the flow and on what Nia should and shouldn't do.
Usability testing changed the design
Buyers left product pages to chat and lost their place. We pivoted from a full-page modal to a persistent sidebar, and wrote the decision back into the spec.
Shipped to production
The persistent sidebar launched. Search and PDP integration moved to the long-term roadmap, also recorded in the spec.
faster purchase flow with Nia
higher bundle attachment
These results belong to the whole Nia team. The workflow's job was to keep every team, and every agent, working from the same page while we got there.
See Nia for yourself
Read how we designed it, or see where it lives in production.
What worked
- The spec never went stale.Docs usually die after kickoff. Because every review ended with a write-back, this one kept up.
- AI got better each week.More decisions in the spec meant better first drafts, without anyone rewriting prompts.
- Hard questions came up early.Empty sections with no owner made the AI-specific unknowns visible at kickoff. confirm
What we're still figuring out
- Spec size.Long specs dilute the context AI actually uses. Next: one module per flow.
- Who breaks a tie.When two teams disagree on a section, the spec needs a clear tie-breaker.
- Measuring quality, not only speed.Faster drafts don't matter if rework goes up.
- Adoption.Teams need a template and a 30-minute setup, not a manifesto.
09 · Rollout
From one pilot to a team standard.
I proposed a staged rollout, so each step earns the next one with evidence instead of enthusiasm.
Pilot on Nia
One AI feature, five teams, one spec, shipped to production.
Proposal and playbook
The spec template, a kickoff checklist, the MCP setup guide, and the guardrails above.
Two more AI features
Two squads start with the template. A short retro every sprint on what the spec missed.
Org standard
No AI feature kicks off without a spec where every section has an owner.
How we'll know it works
Defined before rollout, so the result can't be argued after the fact. Baselines come from the pilot.
| Metric | What it tells us | Baseline | Target |
|---|---|---|---|
| Sections without an owner at kickoff | clarity | TBD | 0 |
| Brief to first design review | speed | TBD | TBD |
| Clarifying questions after handoff | spec quality | TBD | TBD |
| Design system coverage in AI drafts | consistency | TBD | TBD |
| AI features starting with a spec | adoption | 1 | TBD |
10 · Reflection
I didn't need to code to shape how we build AI.
On AI features, the most valuable thing I handed off wasn't a screen. It was clarity: which questions matter, who answers them, and what "right" looks like, written plainly enough for people and machines to act on.
Case study · Insight · AI product process
Living Spec: How Our Product Team Learned to Build AI Features
- Overview
- When our organization started building AI features, no team had done it before. Nobody knew what the model needed, who decided how it should behave, or what should happen when it was wrong. I proposed one living spec that holds everything. Each team claims its part, AI helps draft each one, and every decision is written back until launch.
- Role
- Product Designer · author of the workflowI don't write code. I designed the process from the design and product side: the spec template, the questions every AI feature has to answer, and the review gates. Piloted on Nia, then proposed to the product org.
- Keywords
- AI product processCross-team collaborationModel Context ProtocolAI governance
- Status
- Piloted on Nia · proposed to the product org · still evolving
spec per AI feature, where every team claims its part
of the Figma first pass drafted by AI through MCP, from design system components
phases, each closed by a named human sign-off
change in time from brief to first design review
01 · The problem
Nobody knew how to build an AI feature.
Nia was our first AI feature, and the brief was "let's add AI to the platform." Product, design, engineering, data, and QA each owned a piece. None of us had built something whose answers change, depend on data, and can be wrong. Every team had questions that no one owned.
How do we measure success for something that talks?
How should it behave, and what should it never do?
Which model, and how does it plug into checkout?
Where do the answers come from, and are they current?
How do you test an answer that's different every time?
What is it allowed to say about products and prices?
✕ no shared plan · the answers lived in vendor PDFs, meetings, Slack, and people's heads
No playbook
AI features add questions regular features never had: behavior, data, failure, and evaluation. We had templates for none of them.
Everyone waited on everyone
Each team assumed another owned the hard questions, so they surfaced late, when they were expensive to answer.
Knowledge was scattered
Every new teammate, and every AI prompt, started from zero because the context wasn't written down in one place.
02 · The insight
The missing piece wasn't code. It was one plan every team could read.
I'm a UX designer with a product mindset, and I don't write code. What I can do is turn a messy problem into clear questions and decisions. That turns out to be exactly what AI needs too: plain language, with context. So I designed the workflow around a document, not a tool.
If every question has a section, and every section has an owner, nothing falls through.
One spec holds everything
Evidence, behavior, data, guardrails, decisions. If it isn't in the spec, it doesn't exist.
Every section has an owner
Teams claim their part at kickoff. No question stays unowned.
Plain language first
Anyone can read it, and so can AI. Owning a decision never requires code.
03 · One spec, every team
Every team claims its part of the same file.
Each AI feature starts from the same template. Sections tagged "new for AI" are the questions regular features never had. Pick a team to see what it owns.
Problem & success
What are we solving, and how will we know it worked?
User evidence
What do users actually say and do? Every claim links to a source.
AI behavior
What does it do, how does it sound, and what will it never do?
Data & sources
Where do answers come from, and how fresh are they?
Architecture
Which model, and how does it connect to cart, checkout, and accounts?
Failure & fallback
What happens when it's unsure, wrong, or unavailable?
Guardrails
What is it not allowed to say or do, and who approves the limits?
UI & components
Which design system parts does it use? AI drafts only from these.
Evaluation
Which test prompts must it pass, and what counts as a good answer?
Decision log
What we decided, why, who signed it, and the evidence behind it. Every team writes here.
10 sections, 6 of them new for AI features. Every section has exactly one owner.Product owns 1 section and contributes to 4.Design owns 4 sections and contributes to 4.Engineering owns 1 section and contributes to 5.Data owns 1 section and contributes to 2.QA owns 1 section and contributes to 2.Security & legal owns 1 section and contributes to 1.
04 · How work moves
Five phases. AI drafts, people sign.
The spec starts from two inputs and grows through five phases. Select a phase to see what AI drafts, what a person owns, what the spec gains, and who signs off.
partner docs, platform requirements
research, NPS, tickets, analytics
One file per feature. The single source of truth from research to launch.
phases · each reads from and writes back to the spec
Clusters NPS comments, support tickets, and analytics into themes, each linked to its source.
Product and design validate the themes, pick the problem worth solving, and define success.
§1 Problem & success§2 User evidence
Problem statement agreed before any design starts.
✓ signed by ProductFirst-pass screens and conversation flows in Figma through MCP, built only from design system components. Lists every state, including when the AI is unsure.
How the AI behaves, the flow, the hierarchy, and what it must never do. Deciding what not to build.
§3 AI behavior§6 Fallback§8 Components
Design review with product, engineering, and data.
✓ signed by DesignCoding agents read the same plain-language sections. Drafts tickets straight from the spec.
Engineering owns architecture and model integration, data owns sources. Design answers open questions in the spec, not in Slack.
§4 Data & sources§5 Architecture
Scope confirmed against the spec before the sprint starts.
✓ signed by Eng leadGenerates test prompts and expected answers from the evaluation section, and drafts usability test scripts.
QA judges the answers, not only the click paths. Design runs usability sessions and decides what changes.
§9 Evaluation§10 Decision log
Every test prompt passes, or has a logged exception.
✓ signed by QADrafts release notes and help content, then summarizes post-launch metrics.
Product makes the go / no-go call and security confirms the guardrails. Post-launch data flows back in as new evidence.
§7 Guardrails§2 User evidence
Go / no-go review.
✓ signed by Product + Security↺ every decision is written back into feature-spec.md, until launch
05 · Inside a real spec
Plain language, with an owner on every section.
This is roughly what Nia's spec looked like mid-design. No code anywhere. Engineers, QA, and AI agents all read the same text.
---feature: nia-persistent-sidebarstatus: design · v0.7sources: vendor-input.md, nps (500+ verbatims), usability-r1--- ## Problem @product1Buyers leave the product page to ask Nia a question,then lose their place. ## AI behavior @design2- Suggests products from the buyer's profile- Adds to cart only after the buyer confirms- Never guesses compatibility it can't source ## Data & sources @data- Specs, pricing [src: product catalog]- Buyer context [src: account profile] ## Fallback @design @eng3- Not sure → say so, link to the product page- Checkout issue → hand off to manual checkout ## Decisions @everyone4- v0.6 Modal → persistent sidebar why: keep browsing context while chatting owner: @design · src: usability-r1 · ✓ product ## Evaluation @qa5- "Does this laptop work with our dock?" → cites the spec sheet, or says it isn't sure- Chat history persists across page navigation
An owner on every section
The tag says who answers it. At kickoff, an empty section with no owner is the first red flag.
Behavior, written by design
How the AI acts is a UX decision. Writing it in plain sentences made it something product, engineering, and security could all review.
Failure is designed up front
Every AI feature is wrong sometimes. The fallback section makes "what happens then" a decision, not a surprise.
Decision log
What, why, who, and the evidence. Highlighted: the modal-to-sidebar pivot from usability testing.
Evaluation in questions
QA tests answers, not just click paths. Each test prompt comes with what a good answer looks like.
06 · Designed without code
I write in plain language. AI does the translating.
Every part of the workflow a designer touches is words and structure. AI turns those words into drafts, and a person checks each one. My leverage is the template, the questions it forces, and the judgment at each gate.
AI drafts about 80%. The other 20% is the design job.
AI drafts
- First-pass layouts from design system components
- Every state: empty, loading, error, unsure
- Copy variants and edge-case lists
- Research themes clustered from NPS and tickets
- Tickets and test cases from acceptance criteria
The designer decides
- Which problem is worth solving
- How the AI behaves, and where it stops
- Trade-offs between user needs and business goals
- What not to build, and what moves to long term
- Accessibility, brand fit, and final sign-off
07 · Enterprise guardrails
Speed only counts if legal, security, and brand can trust it.
A process that works for one designer breaks in a Fortune 500 company. These guardrails are what made it something the org could say yes to.
Approved models only
AI runs on enterprise-approved, in-tenant endpoints. Nothing from the spec goes into consumer tools.
Sanitize before merge
Customer PII and partner-confidential terms are stripped from inputs before they enter the spec.
The design system is the fence
Through MCP, AI composes only from published components and tokens. A new component needs design review.
A named owner at every gate
AI never approves anything. Each phase closes with a person signing off inside the spec.
Traceable by default
Every decision records what changed, why, who decided, and the evidence behind it.
Versioned
Every change to the spec has an author and a history, and can be rolled back.
08 · The pilot
We ran it on Nia first.
Nia, our GenAI shopping agent, was the first feature built this way. It was a good test: new territory, a vague brief, and five teams that had never built AI together.
500+ NPS comments, clustered by AI
We fed NPS results from Power BI into Azure OpenAI to find the most common pain points, then validated them with product and engineering in a Lean canvas workshop.
Every team claimed its section
Data and engineering sourced and structured product specs, pricing, and user data so Nia's answers had a real foundation. Design owned behavior and fallback.
First-pass screens drafted through MCP
AI drafted the Figma first pass from our design system. I spent my time on the flow and on what Nia should and shouldn't do.
Usability testing changed the design
Buyers left product pages to chat and lost their place. We pivoted from a full-page modal to a persistent sidebar, and wrote the decision back into the spec.
Shipped to production
The persistent sidebar launched. Search and PDP integration moved to the long-term roadmap, also recorded in the spec.
faster purchase flow with Nia
higher bundle attachment
These results belong to the whole Nia team. The workflow's job was to keep every team, and every agent, working from the same page while we got there.
See Nia for yourself
Read how we designed it, or see where it lives in production.
What worked
- The spec never went stale.Docs usually die after kickoff. Because every review ended with a write-back, this one kept up.
- AI got better each week.More decisions in the spec meant better first drafts, without anyone rewriting prompts.
- Hard questions came up early.Empty sections with no owner made the AI-specific unknowns visible at kickoff. confirm
What we're still figuring out
- Spec size.Long specs dilute the context AI actually uses. Next: one module per flow.
- Who breaks a tie.When two teams disagree on a section, the spec needs a clear tie-breaker.
- Measuring quality, not only speed.Faster drafts don't matter if rework goes up.
- Adoption.Teams need a template and a 30-minute setup, not a manifesto.
09 · Rollout
From one pilot to a team standard.
I proposed a staged rollout, so each step earns the next one with evidence instead of enthusiasm.
Pilot on Nia
One AI feature, five teams, one spec, shipped to production.
Proposal and playbook
The spec template, a kickoff checklist, the MCP setup guide, and the guardrails above.
Two more AI features
Two squads start with the template. A short retro every sprint on what the spec missed.
Org standard
No AI feature kicks off without a spec where every section has an owner.
How we'll know it works
Defined before rollout, so the result can't be argued after the fact. Baselines come from the pilot.
| Metric | What it tells us | Baseline | Target |
|---|---|---|---|
| Sections without an owner at kickoff | clarity | TBD | 0 |
| Brief to first design review | speed | TBD | TBD |
| Clarifying questions after handoff | spec quality | TBD | TBD |
| Design system coverage in AI drafts | consistency | TBD | TBD |
| AI features starting with a spec | adoption | 1 | TBD |
10 · Reflection
I didn't need to code to shape how we build AI.
On AI features, the most valuable thing I handed off wasn't a screen. It was clarity: which questions matter, who answers them, and what "right" looks like, written plainly enough for people and machines to act on.