Skip to main content
Mubienclarity for what comes next
← Selected projects
Build case study

AI Product Launch Review

Turn an AI idea into a launch decision you can test

An interactive product review that turns a system description into explicit claims, boundary tests, evidence requirements, and launch gates. The output is designed to be challenged, checked, and exported.

LiveNext.jsTypeScriptDecision rulesEvaluation design
Try the launch review
AI Product Launch Review screenshot
The problem

Start with the real need

Teams can describe an AI feature as helpful, accurate, or safe without agreeing on what any of those claims would look like in a test. The gap appears later, when a launch review has to connect the product promise to system boundaries, evaluation evidence, and operational ownership.

The approach

How I shaped the product

The prototype asks for the intended use, autonomy, data sensitivity, possible actions, human review model, success condition, and primary concern. Transparent rules then assemble a draft claims register, evaluation pack, evidence checklist, and set of launch gates. The result can be edited and exported as Markdown.

Evidence so farChecked 24 September 2026

What I have actually checked

I ran the transparent rule engine against the three published product profiles: a refund agent, a research copilot, and a shopping agent. The check verified that review levels and capability-specific claims, tests, and gates changed with the stated authority, data, impact, and tools.

3
product profiles checked
19
claims generated
22
evaluation cases generated
20
launch gates generated
  • The refund and shopping agents moved to high scrutiny because they can commit money and create material consequences.
  • The lower-authority research copilot stayed at standard review while still receiving source, untrusted-input, and human-oversight checks.
  • Financial, communication, data, and browsing capabilities each introduced their corresponding claims and adversarial cases.
  • The same inputs produced stable outputs, which makes the current decision logic inspectable and repeatable.
View the worked shopping-agent review →

What this does not establish

This establishes deterministic rule coverage for the public prototype. It does not show that the recommendations are complete, expert-equivalent, or effective in a live product review; those require practitioner comparison and real product evidence.

Product judgment

Decisions that mattered

01

Start with the product claim

A review needs a specific user, task, and consequence. The tool avoids a generic AI risk score and instead asks what the team intends to put into the world.

02

Keep the first engine inspectable

The public prototype uses explicit decision rules rather than an opaque model call. That keeps the relationship between an input and a proposed control visible while the evaluation method is still being tested.

03

Export decisions, not decoration

The useful artifact is a portable review with claims, pass conditions, test cases, owners, and open questions. It should be usable in a repository or product review after the browser tab closes.

Limits

What this does not prove

  • The prototype does not inspect the product or execute the generated evaluation cases.
  • Its output depends on the completeness and accuracy of the system description entered by the user.
  • The suggested controls are decision support, not a safety certification or legal determination.
Next tests

How I would strengthen the evidence

  1. 01Allow teams to attach traces, test results, and review records to each launch claim.
  2. 02Benchmark generated test packs against reviews produced by experienced AI product and risk practitioners.
  3. 03Add specialized packs for agentic commerce, customer support, and multi-agent delegation.

Interested in the reasoning behind the build?

I share the decisions, failed assumptions, and useful lessons as the work develops.

Explore the research