30 Claude Prompts for A/B Tests
Paste a hypothesis, traffic numbers, or raw test results and get back a sized plan, a guardrail set, or a decision-ready read-out your team can use the same day.
In short: This page contains 30 copy-paste ready prompts, organized into 6 categories with a description and pro tip for each. The first 5 prompts are free instantly, no signup needed. Hand-curated and tested by the AI Academy team.
Hypothesis Writing
5 promptsWrite an if/then/because hypothesis from a raw idea
1/30✨ What it does
Turns a one-line test idea into a single if/then/because hypothesis with a clear falsifier.
You are a senior experimentation lead who reviews every test brief before it ships. <context> I have a rough idea for an A/B test and I need a hypothesis my product and marketing leads will accept in one review. </context> <inputs> - Raw idea: [IDEA IN ONE OR TWO SENTENCES] - Page or surface: [PAGE OR FEATURE NAME] - Primary metric: [PRIMARY METRIC] - Audience: [AUDIENCE SEGMENT] - Current baseline: [BASELINE VALUE] </inputs> <task> Turn the idea into a single if/then/because hypothesis, name the smallest change that can falsify it, and list two reasons the idea could fail that are not bad copy. </task> <constraints> Keep the hypothesis to one sentence. Do not invent a lift number. Do not use vague words like engagement or delight. Tie every claim to the stated primary metric. </constraints> <format> Four headers: Hypothesis, What Would Falsify It, Failure Reasons, What I Still Need From You. </format>
Pro tip: Put the real baseline in the last input, or the falsifier section stays abstract and the review will bounce it.
Stress test a hypothesis before it goes to build
2/30✨ What it does
Produces a pre-mortem of a drafted hypothesis so weak assumptions get caught before tickets are written.
You are a skeptical experiment reviewer whose job is to kill weak tests before engineering spends a sprint. <context> I have a drafted hypothesis and I want the holes found before we write tickets. </context> <inputs> - Hypothesis: [IF THEN BECAUSE STATEMENT] - Expected lift: [EXPECTED PERCENT LIFT] - Evidence so far: [RESEARCH, HEATMAPS, SUPPORT TICKETS, OR NONE] - Build cost: [DAYS OF ENGINEERING] </inputs> <task> List the weakest assumptions, one confound that could move the metric with no real user change, and a cheaper pretest that would confirm or kill the idea in under a week. </task> <constraints> Be direct. Cap the assumption list at five. Do not recommend running it longer as the only fix. If the evidence field is NONE, say the hypothesis is not ready and stop there. </constraints> <format> Sections: Weakest Assumptions, Confound To Watch, Cheaper Pretest, Ready To Build (yes or no, one sentence why). </format>
Pro tip: If your evidence field is thin, leave it as NONE on purpose so the prompt blocks the build instead of inventing research.
Convert a qualitative insight into a testable hypothesis
3/30✨ What it does
Pulls one testable hypothesis out of interview or support notes and checks whether traffic can carry it.
You are a conversion researcher who turns interview notes into tests the growth team can actually run. <context> I have qualitative notes from users and I need one testable hypothesis, not a list of themes. </context> <inputs> - Notes: [PASTE INTERVIEW QUOTES, SESSION NOTES, OR SUPPORT TICKETS] - Surface we can change: [PAGE, EMAIL, OR FLOW] - Metric we can measure: [PRIMARY METRIC] - Traffic we can use: [WEEKLY VISITORS ON THAT SURFACE] </inputs> <task> Extract the single strongest pattern in the notes, write one if/then/because hypothesis, and say whether the weekly traffic is enough to test it in four weeks at a 5 percent minimum detectable effect. </task> <constraints> Quote at most two user lines as evidence. Do not invent quotes. If the notes do not support a test, say so and stop. Do not recommend a multivariate test. </constraints> <format> Headers: Pattern Found, Hypothesis, Evidence Quotes, Traffic Verdict. </format>
Pro tip: Paste raw quotes, not your summary. The traffic verdict is only useful if the weekly visitor number is from that same surface.
Write competing hypotheses for the same metric drop
4/30✨ What it does
Writes three mutually exclusive explanations for a metric drop and ranks them by how cheap they are to disprove.
You are a senior analyst who refuses to treat a metric drop as having one cause. <context> Our primary metric dropped and the team already has a favorite explanation. I need competing hypotheses so we do not test the first story we liked. </context> <inputs> - Metric that dropped: [METRIC NAME AND PERCENT CHANGE] - Date the drop started: [DATE] - Favorite explanation: [THE STORY THE TEAM BELIEVES] - What else changed that week: [LAUNCHES, CAMPAIGNS, BUGS, OR UNKNOWN] </inputs> <task> Write three mutually exclusive hypotheses that could explain the drop, rank them by how cheap they are to disprove, and recommend the first check for each. </task> <constraints> Do not pad the favorite explanation. One of the three must treat the drop as a tracking or mix-shift issue, not a product issue. Keep each hypothesis to two sentences. </constraints> <format> A numbered list of three hypotheses. For each: Statement, Disprove Cost (low/med/high), First Check. </format>
Pro tip: Fill the launches and campaigns field even when you think nothing changed. Mix shifts hide in weeks people remember as quiet.
Scope a hypothesis so a two-week test can falsify it
5/30✨ What it does
Cuts a large idea down to one change a two-week test can honestly falsify, or says not to run it.
You are an experiment designer who scopes tests to the shortest window that can still produce a decision. <context> The team wants to test a big idea. I need the hypothesis cut down so a two-week test can actually falsify it. </context> <inputs> - Big idea: [ORIGINAL IDEA] - Two-week traffic: [VISITORS IN TWO WEEKS] - Baseline conversion: [BASELINE CONVERSION RATE] - Decision we need: [SHIP, ITERATE, OR KILL] </inputs> <task> Rewrite the hypothesis so it is about one change on one surface. State the minimum detectable effect a two-week test can support at 80 percent power, and say if that MDE is honest for the decision we need. </task> <constraints> If the two-week traffic cannot support an honest decision, say we should not run the test yet. Do not suggest peeking as a way to finish early. No jargon without a plain-English line next to it. </constraints> <format> Headers: Scoped Hypothesis, Detectable Effect In Two Weeks, Honest Decision (yes or no), What To Cut From The Original Idea. </format>
Pro tip: If the honest-decision line comes back no, shrink the idea again rather than stretching the calendar and peeking.
Sample Size and Power
5 promptsEstimate minimum sample size from baseline and MDE
6/30✨ What it does
Computes per-variant and total sample size from baseline, MDE, and power, then translates it into weeks.
You are a senior experiment statistician who explains sample size to marketers without a stats degree. <context> I need a sample size I can defend in a planning meeting, with the math shown in plain language. </context> <inputs> - Baseline conversion rate: [BASELINE RATE AS A PERCENT] - Minimum detectable effect: [MDE AS A RELATIVE OR ABSOLUTE PERCENT] - Confidence level: [CONFIDENCE LEVEL, USUALLY 95] - Power: [POWER, USUALLY 80] - Traffic split: [SPLIT, E.G. 50/50] - Weekly traffic: [WEEKLY VISITORS] </inputs> <task> Compute per-variant and total sample size. Show the formula in words. Translate the total into weeks using the weekly traffic I gave. </task> <constraints> State whether the MDE is relative or absolute based on how I wrote it. If I mixed them, ask me to pick one and do not invent a number. Round sample size up to the next hundred. </constraints> <format> A short table: Per Variant N, Total N, Weeks. Then four sentences: Formula In Words, Assumptions, Risk If We Stop Early, What Changes If MDE Is Cut In Half. </format>
Pro tip: Say whether your MDE is relative or absolute in the input line. Mixing the two is the usual way this number goes wrong in a planning meeting.
Translate a sample size into a calendar plan
7/30✨ What it does
Turns a target sample size into a week by week calendar with an honest first-read date.
You are a growth operations lead who turns sample size into a weekly calendar the team can staff. <context> I have a target sample size and I need to know when we can start, when we can read, and what would slip the date. </context> <inputs> - Target total sample: [TOTAL USERS OR SESSIONS NEEDED] - Weekly eligible traffic: [WEEKLY ELIGIBLE TRAFFIC] - Ramp plan: [RAMP PERCENT BY WEEK, OR 100 FROM DAY ONE] - Blackout dates: [DATES TO AVOID, OR NONE] - Decision meeting day: [WEEKDAY THE TEAM MEETS] </inputs> <task> Build a week by week plan from start to first valid read. Flag any week where a blackout or ramp makes the date slip. Name the earliest honest read date. </task> <constraints> Do not treat day 7 as a valid read unless the sample is already met. If ramp is slow, say the calendar is longer and show the new date. Keep the plan under 200 words. </constraints> <format> A week table with columns Week, Traffic Captured, Cumulative N, Status. Then one line: Earliest Honest Read Date. </format>
Pro tip: Include sale weeks and product freezes in blackout dates. Those weeks look like traffic and still produce a useless read.
Check if current traffic can support the planned test
8/30✨ What it does
Gives a go or no-go on whether current traffic can power the planned MDE in the weeks you will wait.
You are an experimentation lead who stops underpowered tests before they waste a sprint. <context> We want to run a test on a low-traffic surface and I need a go or no-go based on the numbers, not optimism. </context> <inputs> - Surface: [PAGE OR FLOW] - Weekly traffic: [WEEKLY VISITORS] - Baseline rate: [BASELINE CONVERSION RATE] - Planned MDE: [PLANNED MDE] - Maximum weeks we will wait: [MAX WEEKS] </inputs> <task> Say go or no-go. If no-go, give the MDE we could detect in the max weeks, or the weeks we would need for the planned MDE. Offer one design change that would make a go possible, such as a bigger surface or a binary metric. </task> <constraints> Do not soften a no-go. Do not suggest running at 90/10 if that makes power worse. If weekly traffic is under 200, say this surface is a qualitative test, not an A/B test. </constraints> <format> Verdict on the first line (GO or NO-GO). Then: Math, Detectable MDE In Max Weeks, Design Change That Would Flip The Verdict. </format>
Pro tip: Set max weeks to the real patience of the team, not the number that makes the math work. A six-week no-go is cheaper than a six-week maybe.
Recalculate sample size after a mid-test traffic drop
9/30✨ What it does
Recalculates remaining sample and end date after a traffic drop, and says whether the current lift is usable.
You are an analyst who rescues live tests when traffic drops mid-flight. <context> Our test is live and weekly traffic fell. I need a new end date and a clear call on whether we should stop. </context> <inputs> - Original target N: [ORIGINAL TOTAL SAMPLE] - Users collected so far: [USERS SO FAR] - New weekly traffic: [NEW WEEKLY TRAFFIC] - Days already running: [DAYS LIVE] - Primary metric so far: [OBSERVED LIFT OR NONE] </inputs> <task> Recalculate remaining sample and the new end date. Say whether peeking at the current lift is valid. Recommend continue, stop for futility, or redesign. </task> <constraints> Treat the observed lift as biased if we did not pre-register a sequential rule. Do not let a pretty early lift justify stopping. If remaining time exceeds 4 more weeks, recommend kill or redesign. </constraints> <format> Headers: Remaining N, New End Date, Is Current Lift Usable (yes or no), Recommendation, One Sentence For Slack. </format>
Pro tip: Leave observed lift as NONE if you have been watching the dashboard. The prompt will then refuse to treat that peek as a result.
Size a test that leadership will peek at
10/30✨ What it does
Designs a peeking-safe sequential plan and states the sample size penalty versus a fixed-horizon test.
You are a statistician who designs peeking-safe tests for teams that will look at the dashboard every day. <context> Leadership will open the results daily. I need a plan that does not lie when they do. </context> <inputs> - Baseline rate: [BASELINE CONVERSION RATE] - MDE: [MINIMUM DETECTABLE EFFECT] - How often they will look: [DAILY, EVERY OTHER DAY, OR WEEKLY] - Planned calendar length: [PLANNED WEEKS] - Tool we use: [OPTIMIZELY, VWO, GA4, OR INTERNAL] </inputs> <task> Recommend a sequential or group-sequential plan, the spending function in plain words, and the sample size penalty versus a fixed-horizon test. Give the team a one-line rule for when they are allowed to call the test. </task> <constraints> Name the tool limitation if the stated tool cannot do sequential testing. Do not invent p-values. Keep the explanation free of Greek letters. </constraints> <format> Sections: Recommended Plan, Sample Size Penalty, Call Rule, Tool Limitation. </format>
Pro tip: If your tool is GA4, say so. The tool-limitation section will stop you from pretending you have sequential testing when you do not.
Guardrail Metrics
5 promptsPick guardrail metrics for a conversion test
11/30✨ What it does
Chooses three ship-blocking guardrails (harm, revenue quality, ops) tied to events you can already query.
You are a senior product analyst who chooses guardrails so a conversion win cannot hide damage. <context> We are testing a change meant to lift conversion and I need guardrails that would stop a ship even if the primary wins. </context> <inputs> - Primary metric: [PRIMARY METRIC] - Surface: [PAGE OR FLOW] - Business model: [SUBSCRIPTION, ECOMMERCE, LEAD GEN, OR OTHER] - Known risks: [RISKS, E.G. DISCOUNT ABUSE OR SUPPORT LOAD] - Data we can query daily: [AVAILABLE EVENTS] </inputs> <task> Recommend three guardrail metrics: one user harm, one revenue quality, one operational. For each, set a ship-blocking threshold and a watch threshold. </task> <constraints> Every guardrail must be measurable with the events listed. If an event is missing, name the event to instrument instead of faking a metric. Do not use vanity metrics as guardrails. </constraints> <format> A table: Guardrail, Type, Ship-Blocking Threshold, Watch Threshold, Event Needed. </format>
Pro tip: List only events that exist today. Missing events should come back as instrumentation work, not as fake guardrails.
Set kill criteria for a guardrail breach
12/30✨ What it does
Writes a one-query kill playbook with actor, rollback steps, and a ready channel message.
You are an experiment operations lead who writes kill criteria the on-call person can follow at 11pm. <context> I need a kill rule for a live test so nobody has to debate in Slack when a guardrail breaks. </context> <inputs> - Guardrail metric: [GUARDRAIL METRIC] - Baseline value: [BASELINE VALUE] - Breach we will not accept: [BREACH THRESHOLD] - Who can kill the test: [ROLE OR NAME] - Rollback method: [FEATURE FLAG, REVERT, OR CONFIG] </inputs> <task> Write a kill playbook: the exact condition, the confirmation check, who acts, how to roll back, and the message to post in the experiment channel. </task> <constraints> The condition must be checkable in one query. No use-judgment lines. If the rollback method is CONFIG, list the two fields to change. Keep the Slack message under 80 words. </constraints> <format> Headers: Kill Condition, Confirm In One Query, Actor, Rollback Steps, Channel Message. </format>
Pro tip: Name a single person or role as the killer. Shared ownership is how a breached guardrail stays live overnight.
Diagnose a winning primary with a failing guardrail
13/30✨ What it does
Gives a one-word ship, iterate, or kill when the primary wins and a guardrail fails.
You are a principal analyst who decides what to do when the primary wins and a guardrail fails. <context> The test looks like a win on the primary metric and a loss on a guardrail. I need a recommendation, not a hedged paragraph. </context> <inputs> - Primary result: [PRIMARY METRIC, LIFT, AND CONFIDENCE] - Guardrail result: [GUARDRAIL METRIC, LIFT, AND CONFIDENCE] - Test length and N: [DAYS AND SAMPLE SIZE] - Who is pressuring to ship: [TEAM OR STAKEHOLDER] - Revenue impact if we ship anyway: [ROUGH REVENUE NOTE] </inputs> <task> Recommend ship, iterate, or kill. Explain the trade in numbers. Propose one follow-up that would resolve the conflict in a smaller, safer test. </task> <constraints> Do not recommend shipping a failed revenue or retention guardrail to please a stakeholder. If sample is thin, say the conflict is not even real yet. Keep the recommendation to one word on the first line. </constraints> <format> First line: SHIP, ITERATE, or KILL. Then: Trade In Numbers, Why The Pressure Is Wrong Or Right, Follow-Up Test. </format>
Pro tip: Paste the raw interval, not a rounded lift. A guardrail that includes zero is a watch, not an automatic kill.
Design a revenue guardrail for a UX or copy test
14/30✨ What it does
Defines a paired revenue guardrail so a conversion win cannot ship if order value or plan mix falls too far.
You are a monetization analyst who stops UX tests from buying conversion with weaker revenue. <context> We are testing copy or layout. Conversion may rise while average order value or plan mix falls. I need a revenue guardrail that catches that. </context> <inputs> - What we are testing: [COPY, LAYOUT, OR UX CHANGE] - Conversion metric: [CONVERSION METRIC] - Revenue metric available: [AOV, ARPU, LTV PROXY, OR PLAN MIX] - Current values: [CURRENT CONVERSION AND REVENUE VALUE] - Acceptable revenue dip: [MAX PERCENT DIP] </inputs> <task> Define a composite or paired guardrail so a conversion win with a revenue dip past the limit cannot ship. Show a numeric example using the current values. </task> <constraints> Do not use revenue per visitor if I did not list it as available. Show the example with the numbers I gave. Flag if the acceptable dip is so wide the guardrail is decorative. </constraints> <format> Headers: Guardrail Definition, Numeric Example, Ship Rule, Is The Dip Limit Honest. </format>
Pro tip: Set the dip limit from last quarter's normal week-to-week swing, not from a number that feels comfortable in the meeting.
Write a guardrail dashboard brief for analytics
15/30✨ What it does
Writes a one-page analytics brief for a daily guardrail dashboard with day-one segment rules.
You are a product analytics lead who briefs the person building the experiment dashboard. <context> I need a one-page brief so analytics can ship a daily guardrail view before the test starts, not after. </context> <inputs> - Test name: [TEST NAME] - Primary metric: [PRIMARY METRIC AND EVENT] - Guardrails: [LIST OF GUARDRAIL METRICS AND EVENTS] - Segment cuts we allow: [ALLOWED SEGMENTS] - Tool: [LOOKER, MODE, AMPLITUDE, OR OTHER] </inputs> <task> Write a dashboard brief: tiles, default date range, the one segment cut that is allowed on day one, and the tiles that must stay hidden until sample is met. </task> <constraints> Forbid more than one segment cut on day one. Name the exact events. If the tool is AMPLITUDE, specify a chart type that tool can actually do. No nice-to-have section. </constraints> <format> Headers: Tiles In Order, Default Range, Allowed Cut, Hidden Until N Met, Event Map. </format>
Pro tip: Send this brief before the flag flips. A dashboard built after launch is how people peek on cuts you never registered.
These prompts give you the what. Tutorials give you the why.
Learn when to use extended thinking, how to build Claude Projects, and workflows that compound. 300+ tutorials and growing.
Variant and Assignment Design
5 promptsDesign control and variant from a hypothesis
16/30✨ What it does
Turns a hypothesis into a ticket-ready control and variant spec with QA checks and an isolation rule.
You are a senior CRO designer who turns a hypothesis into two variants a developer can build without a meeting. <context> I have a hypothesis and I need a control and one variant specified tightly enough to ticket. </context> <inputs> - Hypothesis: [IF THEN BECAUSE STATEMENT] - Current control: [WHAT USERS SEE TODAY] - Brand or legal limits: [WORDS OR CLAIMS WE CANNOT USE] - Device mix: [DESKTOP PERCENT / MOBILE PERCENT] - Designer available: [YES OR NO] </inputs> <task> Specify control and variant: copy, layout, CTA, and what must stay identical so the test isolates one idea. List acceptance checks for QA. </task> <constraints> Change only what the hypothesis needs. If designer is NO, keep the variant to copy and a simple layout shift a developer can do. Respect legal limits exactly. </constraints> <format> Headers: Isolation Rule, Control Spec, Variant Spec, QA Checks. </format>
Pro tip: If designer is NO, say so. The spec will stay in copy and spacing, which is what a developer can ship in a day.
Write a multivariate plan with isolation rules
17/30✨ What it does
Chooses A/B, A/B/C, or sequenced tests so a bundle of changes does not wipe out the learning.
You are an experiment architect who stops teams from testing three ideas in one messy variant. <context> The team wants to change headline, image, and CTA at once. I need a plan that isolates learning. </context> <inputs> - Changes they want: [LIST OF CHANGES] - Weekly traffic: [WEEKLY TRAFFIC] - Primary metric: [PRIMARY METRIC] - Time we have: [WEEKS AVAILABLE] - Constraint from brand: [BRAND RULE] </inputs> <task> Recommend A/B, A/B/C, or a sequenced pair of tests. If multivariate is possible, say so with the cell count and the traffic each cell gets. State what we will not learn if we bundle. </task> <constraints> If weekly traffic cannot support more than two cells, forbid multivariate. Do not use a full factorial as a default. Keep the recommendation to one design. </constraints> <format> First line: Recommended Design. Then: Why, What We Will Learn, What We Will Not Learn, Ticket Breakdown. </format>
Pro tip: Paste the full wish list. The what-we-will-not-learn section is the part that stops a three-idea variant from shipping as one cell.
Split traffic without biasing assignment
18/30✨ What it does
Recommends an assignment unit, persistence, and cache rules so users do not flicker between variants.
You are a senior experiment engineer who reviews assignment plans for bias. <context> I need an assignment plan that does not leak, flicker, or put the same user in two variants across devices. </context> <inputs> - Assignment unit: [USER ID, COOKIE, ACCOUNT, OR SESSION] - Logged-in share: [PERCENT LOGGED IN] - Cross-device concern: [YES OR NO] - Edge cache: [CDN OR NONE] - Tool: [FLAG SYSTEM OR TESTING TOOL] </inputs> <task> Recommend the assignment unit, how to persist it, how to avoid cache flicker, and the QA checks that prove a user cannot see both variants in one session. </task> <constraints> If the unit is SESSION, warn that returning users will switch and the metric will be noisy. If edge cache is CDN, require a cache key that includes the variant. No marketing language. </constraints> <format> Headers: Unit, Persistence, Cache Rule, Flicker Risk, QA Proof. </format>
Pro tip: If you have a CDN, say CDN. A missing cache key is the usual reason a home-page test looks like a win and a mess at the same time.
Design a holdout for an always-on personalization
19/30✨ What it does
Sizes a durable holdout and states when it can shrink so always-on personalization stays measurable.
You are a measurement lead who keeps a holdout so personalization does not become an unmeasured habit. <context> We want to leave a treatment on after the test. I need a holdout so we can still measure incrementality next quarter. </context> <inputs> - Treatment: [FEATURE OR PERSONALIZATION] - Eligible population: [ELIGIBLE USERS] - Metric: [METRIC NAME] - Holdout we can politically defend: [MAX HOLDOUT PERCENT] - Review cadence: [MONTHLY OR QUARTERLY] </inputs> <task> Recommend holdout size, assignment, the incrementality calculation, and the rule for when the holdout can shrink. </task> <constraints> If the political max holdout is under 5 percent and the population is small, say incrementality will be too noisy and recommend a periodic pause instead. Do not promise a weekly read on a 5 percent holdout. </constraints> <format> Headers: Holdout Size, Assignment, Incrementality Math, Shrink Rule, Risk Flag. </format>
Pro tip: Write the political max as the real number you can defend, not 10 percent. The risk flag is there to stop a decorative holdout.
Plan a test that must read on mobile and desktop separately
20/30✨ What it does
Plans a device-split read so a desktop win cannot hide a mobile loss, with a ship rule when they disagree.
You are a CRO lead who has been burned by a desktop win that hid a mobile loss. <context> Device mix is uneven and I will not ship a blended win. I need a plan that reads mobile and desktop as separate decisions. </context> <inputs> - Desktop weekly traffic: [DESKTOP WEEKLY VISITORS] - Mobile weekly traffic: [MOBILE WEEKLY VISITORS] - Baseline by device: [DESKTOP RATE / MOBILE RATE] - MDE we care about: [MDE] - What we can change per device: [SHARED OR SEPARATE LAYOUTS] </inputs> <task> Say whether to run one test with a pre-registered device cut or two tests. Size each read. State the ship rule if one device wins and the other loses. </task> <constraints> Pre-register the device cut. Do not allow a blended primary to override a mobile loss. If one device lacks traffic, say that device is qualitative only. </constraints> <format> Headers: Design Choice, Desktop N And Weeks, Mobile N And Weeks, Ship Rule When Devices Disagree. </format>
Pro tip: If layouts are shared, say SHARED. The ship rule will then force a follow-up rather than a blended rollout that hurts phones.
Experiment Read-outs
5 promptsWrite a decision-ready experiment read-out
21/30✨ What it does
Writes a one-page read-out that opens with a decision and only uses pre-registered segments.
You are a senior experimentation lead who writes read-outs that end in a decision, not a chart dump. <context> The test has reached its pre-registered sample and I need a one-page read-out for the decision meeting. </context> <inputs> - Hypothesis: [ORIGINAL HYPOTHESIS] - Primary result: [PRIMARY METRIC, LIFT, INTERVAL, P VALUE OR POSTERIOR] - Guardrails: [EACH GUARDRAIL AND RESULT] - Sample and duration: [N AND DAYS] - Segments we pre-registered: [SEGMENTS OR NONE] </inputs> <task> Write a read-out that states the decision, the primary result in one sentence, guardrail status, and the one caveat that could change the decision if it is wrong. </task> <constraints> The first sentence must be the decision. Do not add segments that were not pre-registered. If the interval includes zero, do not call it a win. Keep the page under 350 words. </constraints> <format> Headers: Decision, Primary, Guardrails, Caveat, What Happens Next Week. </format>
Pro tip: If you did not pre-register segments, put NONE. Extra cuts belong in a follow-up prompt, not in this page.
Explain a non-significant result to leadership
22/30✨ What it does
Turns a non-significant result into a six-beat script that states what you can rule out and the next cheaper action.
You are a product analytics lead who explains a flat result without making it sound like a waste. <context> The test did not reach significance and leadership is asking what we got for the three weeks. </context> <inputs> - Hypothesis: [HYPOTHESIS] - Result: [LIFT AND INTERVAL] - Sample vs plan: [PLANNED N VS ACTUAL N] - What we observed qualitatively: [NOTES] - Cost of the test: [ENGINEERING DAYS AND LOST CONVERSION IF ANY] </inputs> <task> Write a 6-beat verbal script: what we asked, what the interval allows us to rule out, what we still do not know, and the next action that is cheaper than a rerun. </task> <constraints> Do not call a non-significant lift directional unless the interval excludes a harmful effect we care about. Do not recommend an immediate rerun of the same test. Keep each beat to four spoken sentences. </constraints> <format> Six labeled beats: Ask, Result, What We Can Rule Out, What We Cannot Claim, Cost, Next Cheaper Action. </format>
Pro tip: Paste the interval, not just the p-value. Ruling out a 10 percent loss is often the thing leadership actually needed.
Segment a result without fishing
23/30✨ What it does
Allows only pre-registered cuts and parks extra slices as new tests instead of invented wins.
You are a careful analyst who only cuts data that was written down before the test launched. <context> People want to slice the result by country, device, and plan. I need a rule-bound read so we do not invent a win in a cell. </context> <inputs> - Pre-registered segments: [LIST, OR NONE] - Requested extra cuts: [LIST PEOPLE WANT] - Cell sizes: [N PER REQUESTED CUT IF KNOWN] - Primary result overall: [OVERALL LIFT AND INTERVAL] - Decision at stake: [SHIP OR ITERATE] </inputs> <task> Allow only the pre-registered cuts. For each requested extra cut, say look, ignore, or park for a new test. If a cell is small, say the number that would make the cut honest. </task> <constraints> If pre-registered segments is NONE, the only honest read is the overall result. Do not let a pretty cell override a flat overall. No interesting language. </constraints> <format> A table: Cut, Status (pre-registered / extra), Verdict (use / ignore / new test), Note. </format>
Pro tip: Keep a dated note of pre-registered cuts in the experiment brief. If that note does not exist, this prompt should return overall only.
Detect novelty and seasonality in a result
24/30✨ What it does
Judges whether an early lift is novelty, season, or stable, and says how many more days to run.
You are an experiment analyst who checks whether a lift is novelty, season, or a real change. <context> The lift showed up fast and I do not trust it. I need a check list before we call a win. </context> <inputs> - Daily primary series: [PASTE DATE, CONTROL RATE, VARIANT RATE] - Launch date: [LAUNCH DATE] - Known season: [SALE, HOLIDAY, PAYDAY, OR NONE] - Variant type: [COPY, UI, PRICING, OR ALGORITHM] - How long novelty lasted last time: [DAYS OR UNKNOWN] </inputs> <task> Judge novelty vs season vs stable lift. Point to the days that drive the call. Say how many more days we should run if the call is unstable. </task> <constraints> If the series is shorter than 14 days, say you cannot separate novelty from a real lift. Do not average away a Monday spike. If variant type is PRICING, treat season as the default suspect. </constraints> <format> Headers: Call (NOVELTY, SEASON, STABLE, OR TOO SHORT), Days That Matter, Extra Days To Run, One Chart Title To Plot. </format>
Pro tip: Paste daily rates, not a weekly average. A Monday spike disappears in the weekly number and that is usually the whole story.
Compare two tests that overlapped on the same users
25/30✨ What it does
States whether two overlapping tests can be read independently and what to tell both teams.
You are a senior analyst who untangles overlapping experiments. <context> Two tests ran on overlapping traffic and both teams want credit. I need a clean statement of what we can and cannot attribute. </context> <inputs> - Test A: [NAME, HYPOTHESIS, RESULT] - Test B: [NAME, HYPOTHESIS, RESULT] - Overlap: [PERCENT OF USERS IN BOTH] - Assignment: [INDEPENDENT, MUTEX, OR UNKNOWN] - Shared metric: [SHARED METRIC] </inputs> <task> State whether the results can be read independently. If not, describe the interaction risk and the analysis that would separate them, or say we cannot separate them with the data we have. </task> <constraints> If assignment is UNKNOWN, treat interaction as likely. Do not average the two lifts. If overlap is over 30 percent, say neither read-out is clean. </constraints> <format> Headers: Can We Read Independently, Interaction Risk, What To Compute, What We Should Tell Both Teams. </format>
Pro tip: If you do not know the assignment method, put UNKNOWN. Pretending the flags were mutex is how both teams ship a fake win.
Most people use 10% of Claude. Tutorials unlock the rest.
AI Academy: 300+ hands-on tutorials on Claude, ChatGPT, Midjourney, and 50+ AI tools. New tutorials added every week.
Ship, Iterate, or Kill
5 promptsRecommend ship, iterate, or kill
26/30✨ What it does
Returns a one-word ship, iterate, or kill call with the two numbers and the sentence to say in the room.
You are a head of growth who makes ship, iterate, or kill calls in a 20-minute review. <context> I have a finished test and I need a one-word recommendation I can defend to product, design, and finance. </context> <inputs> - Read-out summary: [PASTE THE DECISION PAGE OR KEY NUMBERS] - Guardrail status: [PASS, WATCH, OR FAIL] - Cost to ship: [ENGINEERING AND SUPPORT COST] - Cost to iterate: [DAYS TO A FOLLOW-UP] - Political pressure: [WHO WANTS WHAT] </inputs> <task> Recommend SHIP, ITERATE, or KILL. Give the two numbers that justify it. Write the sentence I should say out loud in the meeting. </task> <constraints> A failed guardrail cannot be a ship. Political pressure cannot flip the call. If the read-out is incomplete, return BLOCKED and list the missing number. </constraints> <format> First line: SHIP, ITERATE, KILL, or BLOCKED. Then: Two Numbers, Spoken Sentence, What The Losing Side Will Argue, Your Reply. </format>
Pro tip: If finance is in the room, put cost to ship in money, not days. The two-numbers section is what stops the debate from going circular.
Write a follow-up test from a winning variant
27/30✨ What it does
Picks one follow-up from a win, keeps the winner as control, and says whether a holdout is required first.
You are an experiment lead who turns a win into the next isolated question instead of a pile of extra changes. <context> Variant B won. The team wants to add three more ideas on top. I need the next test that learns one thing. </context> <inputs> - Winning variant: [WHAT WON AND THE LIFT] - Ideas people want to pile on: [LIST OF IDEAS] - Remaining traffic: [WEEKLY TRAFFIC AFTER ROLLOUT] - Risk if we bundle: [RISK NOTE] - Time to next review: [WEEKS] </inputs> <task> Pick one follow-up hypothesis from the list. Explain why the others wait. Spec control (the winner) and the next variant. </task> <constraints> Control must be the winning variant, not the old control. Only one new change. If remaining traffic is low because of a full rollout, say we need a holdout first. </constraints> <format> Headers: Next Hypothesis, Why The Others Wait, Control, Next Variant, Holdout Needed (yes or no). </format>
Pro tip: After a full rollout, remaining traffic is often near zero. Say so, or the prompt will design a test you can no longer assign.
Document a failed test as a reusable learning
28/30✨ What it does
Writes a short wiki entry for a lost or flat test, including search phrases later teams will type.
You are the owner of the experiment wiki and you write entries people still use a year later. <context> The test lost or was flat. I need a wiki entry so nobody reruns the same idea with new adjectives. </context> <inputs> - Hypothesis: [HYPOTHESIS] - Result: [RESULT WITH NUMBERS] - What we think happened: [WORKING THEORY] - Artifacts: [LINKS OR FILE NAMES] - Who will search for this later: [TEAM NAMES] </inputs> <task> Write a wiki entry: what we tested, what happened, the theory, the search terms those teams will use, and the one idea that is still allowed. </task> <constraints> Do not reframe a loss as a branding win. Keep it under 250 words. Include three search phrases the later reader will actually type. </constraints> <format> Headers: What We Tested, What Happened, Working Theory, Search Phrases, Still Allowed. </format>
Pro tip: Name the teams who will search, not just growth. Design and support rerun old tests when the wiki only uses experiment-lead language.
Brief engineering on a winning rollout
29/30✨ What it does
Writes a rollout brief with ramp steps, a 50 percent hold, events that must keep firing, and a named kill switch.
You are a product manager who writes the rollout ticket after a win so engineering does not re-litigate the test. <context> We decided to ship the winner. I need a brief that covers ramp, flags, analytics, and the kill switch. </context> <inputs> - Winning variant: [VARIANT SPEC] - Flag name: [FEATURE FLAG] - Ramp plan: [PERCENT BY DAY] - Analytics events that must keep firing: [EVENT NAMES] - Kill switch owner: [ROLE] </inputs> <task> Write an engineering brief: what to make default, ramp steps, events that must not break, and the condition that pauses the ramp. </task> <constraints> Do not ask engineering to clean up the experiment code in this ticket. Ramp must include a 24-hour hold at 50 percent. Name the kill switch owner in the first paragraph. </constraints> <format> Headers: Default State, Ramp Steps, Events That Must Keep Firing, Pause Condition, Out Of Scope. </format>
Pro tip: Put the flag name exactly as it exists in the system. A renamed flag in the ticket is how the kill switch points at nothing.
Prepare a weekly experiment review agenda
30/30✨ What it does
Builds a 30-minute review agenda that decides finished tests first and ends with a queue cut.
You are a director of experimentation who runs a 30-minute weekly review that actually decides things. <context> I have a messy list of live, queued, and finished tests. I need a 30-minute agenda with time boxes and a decision for each finished test. </context> <inputs> - Live tests: [LIST WITH DAYS LIVE AND CURRENT N] - Queued tests: [LIST WITH OWNERS] - Finished this week: [LIST WITH DRAFT DECISIONS] - Team in the room: [ROLES] - Recurring blocker: [BLOCKER] </inputs> <task> Build a 30-minute agenda. Put finished tests first. Give each live test one question, not a tour. End with a queue cut that fits capacity. </task> <constraints> 30 minutes total. No status tours. If more than three finished tests, take the two that need a fight and park the rest in writing. Name the person who speaks, not the team. </constraints> <format> A timed agenda table: Minutes, Item, Speaker, Decision Needed. Then a one-line queue cut. </format>
Pro tip: Name roles that will be in the room, not the full org chart. The speaker column only works if those people are actually attending.
Free tool
Prompt Optimizer
Turn a rough idea into a structured, professional AI prompt.
Frequently Asked Questions
Prompts are the starting line. Tutorials are the finish.
A growing library of 300+ hands-on tutorials on ChatGPT, Claude, Midjourney, and 50+ AI tools. New tutorials added every week.
7-day free trial. Cancel anytime.
Related guides