The GTM Experiment Playbook

How to find out whether a segment, an offer or a feature idea gets a response from the market, with a few hundred emails: hypothesis, target rate, delivery gate, sample size, stopping rules, and how to read the verdict.

Updated September 15, 2026

This playbook is for B2B teams that don't yet know who to sell to, or what to sell them. It turns "let's try cold email" into an experiment: one question, a target decided in advance, a sample size decided in advance, and a yes or no you can defend.

It is written for two readers at once. A person can read it top to bottom. An AI agent connected to the Sellvance MCP server can follow it step by step: every rule is stated as a threshold or a check, and the appendix lists the tool calls in order.

The questions an experiment answers:

  • Does this ICP respond, or not?
  • Does this offer work, or not?
  • Does our ICP need this feature, or not?

What an experiment gives you: the answer, including "no", in weeks rather than months.

What it does not give you: leads, meetings, or revenue. Those can happen along the way. They are not what the experiment is for.


1. What an answer costs

Each experiment tests one campaign against a threshold: the reply rate at which the answer is "yes, this works". The Signal Check in Sellvance tells you when the result is conclusively above it, or conclusively below it.

This playbook never tells you what reply rate is normal. Published "average reply rates" come from the customers of the tools that publish them, and they mix products, markets and offers that have nothing in common. What it tells you instead is what it costs to find out:

What's tested Threshold If the true rate is Sends to a verdict
Delivery (zero out-of-office replies) 1% 0% 299
Interest 3% 6% ~207
Interest 3% 8% ~83
Interest 3% 10% ~42
Interest 3% 15% 30
Interest 5% 15% ~36
Interest 10% 25% ~31

The delivery row is exact: after 299 sends with no auto-reply at all, a 1% baseline is rejected. The interest rows are the sends for an 80% chance of a "works" verdict, computed from the Signal Check's own test (exact binomial, 90% Clopper-Pearson interval). 30 is the smallest sample the Signal Check will judge at all.

None of these rows says what reply rate you will get. Each says how many emails it takes to know. The stronger your offer is against your threshold, the cheaper it is to prove. A precise segment and a message written for it are what move you down the table.

If the true rate sits right at your target, no sample size will give a clean verdict. That is a real answer too (section 6, grey zone).


Gate: is cold outreach the right channel?

Cold outreach doesn't work for every ICP. Check these before you design an experiment. They are a hard gate, not advice: if any one of them is true for your ICP, don't design the experiment. The answer is already "no", and no amount of copy changes it.

Condition Why the channel fails
The ACV can't pay for the channel A self-serve product at $20 a month can't pay for domains, mailboxes and the time spent handling replies. Check it with the break-even threshold in section 2: if it's above the rate you would honestly act on, stop.
The buyer doesn't read work email Micro-business owners, developers, doctors, tradespeople. Silence tells you about their inbox, not about your offer.
The persona is saturated For example a VP of Sales or a CTO at a funded startup: filters, assistants and a learned reflex bury cold email before anyone reads it.
There is no trigger If nothing makes the problem urgent at a particular moment, the email always arrives at the wrong time.
The jurisdiction requires prior consent Where the law requires opt-in before contact, the sample you can legally reach is zero.
The value can't be understood without a demo If the product can't be explained in about 80 words, replies say "didn't get it", not "no thanks". You'd be testing the explanation, not the offer.

If one of these applies, record which one and stop. Finding that out in an afternoon is the cheapest answer this playbook can give.


2. Write the hypothesis

One experiment answers one question. Write it down before you build anything:

Segment × Persona × Offer → reply rate of at least Threshold

  • Segment: the companies. Narrow enough that the pain is the same across all of them. Not "B2B SaaS", but "bootstrapped B2B SaaS, 2–20 people, no sales hire".
  • Persona: the role you email at those companies. Founder, head of sales, operations lead.
  • Offer: the one promise the email makes, and the yes/no action it asks for.
  • Threshold: the reply rate below which you would not act on this. Decide it before you see any data, and never move it afterwards.
  • Expected rate: the rate you want to be able to prove. It sizes the experiment (section 5). It is an assumption for sizing, not a forecast.

Where the threshold comes from

Not from an industry average. The question isn't "what reply rate is normal?" but "what reply rate makes this channel pay for itself for us?". Cold email pays off when

ACV × reply-to-deal conversion × replies > the cost of the sends and your time

so the break-even reply rate per email sent is

threshold = all-in cost per email ÷ (ACV × reply-to-deal conversion)

A product with a $30,000 ACV where one reply in ten closes breaks even at 1.5% if each email costs $45 all-in, counting your time. A $600-a-year product with the same conversion needs 12% even at $7.20 per email. There is no single right number; there is only yours. The experiment planner calculates it for you.

What each kind of question holds fixed:

Question What the experiment tests Example hypothesis
Does this ICP respond? The segment Bootstrapped B2B SaaS founders reply to offer X at ≥ 10%
Does this offer work? The promise in the email Agency owners reply at ≥ 8% to "we find your next five clients"
Does our ICP need this feature? The problem the feature solves, offered before it is built Ops leads at 50–200 person logistics firms reply at ≥ 10% to "stop re-keying orders from email"

One experiment changes one thing. If "doesn't work" comes back, change a single element (the segment, the persona, or the offer) and run a new experiment on fresh leads. Changing two at once leaves you unable to say which one mattered.

Rates are per email sent

The Signal Check counts every email sent, across all steps of the sequence. In a two-step sequence, everyone who doesn't reply to the first email gets a second, so the rate per email is roughly half the rate per lead. Set the target in the same unit the Signal Check uses: replies per email sent.


3. Gate: confirm delivery before you trust any result

A reply rate mixes two things: whether the email arrived, and whether it persuaded. Zero replies is equally consistent with a bad offer and with a spam folder, and the two need opposite fixes. A "doesn't work" verdict is only valid once delivery is confirmed. This is a hard gate, not a recommendation.

Before sending:

  1. Every sending domain has SPF, DKIM, DMARC and MX published and verified.
  2. Every sending mailbox passes its health check and has IMAP connected. Without IMAP, replies are never counted.
  3. Run an inbox placement test on the real rendered first message, from each sending mailbox. Caveat: the test lands in one of your own IMAP mailboxes, so it catches authentication failures and spam-filter scores, but "inbox" there is not proof of inboxing at Gmail or Outlook.

While sending:

  1. Read the delivery signal on the campaign's Signal Check first. Delivery is measured with out-of-office auto-replies: mail servers send them without a human reading anything, and Gmail and Outlook suppress them for mail classified as spam. So the auto-reply rate reflects inbox placement and nothing else.
  2. If delivery fails, interest is reported as blocked. Do not read the reply rate. Pause, fix infrastructure, then start again on fresh leads.
  3. Zero auto-replies has two possible causes: spam placement, or replies not being collected (no IMAP connection). Rule out the second before concluding the first.

Auto-reply rates are seasonal: roughly 1–3% in ordinary weeks, 8–12% in August and late December. The delivery threshold uses your workspace's own measured rate once it has at least 200 sends of history.


4. Design

4.1 A fresh sample

Every lead in the experiment must be a company nobody from your side has contacted. A company that already got an email from you is answering a second email, not your offer.

The database holds 9.5 million companies, so fresh leads are not the constraint. Keeping them fresh is. Sellvance does it with the saved search's rejected list: a saved search never adds a company that has any lead in that segment.

  1. Check that the workspace has an All leads segment: a segment with one filter that has no conditions. New workspaces come with one. Create it if it's missing.
  2. Save that segment once before the experiment starts. Saving re-evaluates its filter, so every lead already in the workspace becomes a member. From then on, new leads join it as they're added.
  3. Set the rejected list of the experiment's saved search to All leads.

Every company that an earlier experiment or campaign touched is now excluded.

4.2 One target per workspace

The target reply rate is a workspace setting. Experiments that share a target can run side by side in one workspace. An experiment with a different target needs its own workspace, or has to wait until the others finish.

Experiments running side by side are each judged against the target. Don't compare them with each other: two campaigns that both pass are both "yes", not a ranking.

4.3 Hold everything fixed while it runs

Once the first email of the experiment is sent, don't change:

  • the sequence, its steps, subjects or copy
  • the sending mailboxes
  • the target reply rate
  • the segment's filters or enrichment chain

Any of these changes starts a new experiment. The sends before the change don't count toward it.


5. Size and run the experiment

  1. Pick the planned sample. Enter your threshold and the rate you want to be able to prove into the experiment planner, or read it off the table in section 1. When unsure, size for the lowest rate you would still act on: it takes more sends, but the verdict holds either way. Write the number down: it's the point where you read the verdict.
  2. Validate by hand first. Build the campaign with autopilot off, add 3–5 leads manually, and check everything end to end: the enrichment filled every field, the generated emails read well, the placement test lands in the inbox. Manual adding exists for exactly this check.
  3. Turn on autopilot. Once validation passes, enable autopilot and set the campaign's daily limit (add_limit). From then on, autopilot pulls new companies through the saved search, runs the enrichment chain, enrols qualified leads, and generates and reviews the emails.
  4. Stop enrolling at the planned sample. Leads needed ≈ planned sample ÷ number of steps. When that many leads are enrolled, set add_limit to 0. Follow-ups to leads already enrolled keep going.
  5. Read the verdict once the Signal Check shows the planned number of sends and the last follow-ups have gone out (section 6).

If the planned sample is more than you can afford, change the experiment, not the rule: test a narrower segment where you expect a higher reply rate, or a bolder offer.


6. Stopping rules, decided in advance

Write these down next to the hypothesis before the first email goes out.

Outcome Signal Check at the planned sample What to do
Works Delivery passes, interest passes The segment × persona × offer beats the target. Scale it: start a separate campaign with a bigger daily limit and more mailboxes. Nothing on the platform stops you. Keeping it separate leaves the experiment's numbers clean.
Doesn't work Delivery passes, interest fails This combination is conclusively below the target. Change one element and run a new experiment on fresh leads.
Grey zone Delivery passes, interest still inconclusive The true rate is close to the target. Extend once, by at most 50% of the planned sample. If it is still inconclusive, the answer is "not clearly above target": treat it as no, and move on.
Harmful Harm fails at any point Unsubscribes are conclusively above your limit. Pause the campaign. The offer or the list is wrong for these people, whatever the reply rate says.
Invalid Delivery fails, or interest is blocked No conclusion about the market is possible. Pause, fix delivery, restart on fresh leads.

Delivery and harm can end an experiment early. Interest is read only at the planned sample (section 7.2 explains why).


7. Reading the verdict

7.1 The Signal Check

Every campaign, and the workspace as a whole, has a Signal Check with three signals, each an exact binomial test against a threshold:

Signal Counts Fails when
Delivery Out-of-office auto-replies Too few
Interest Replies written by a person Too few
Harm Unsubscribes Too many

Each signal has one of four statuses:

  • pass: conclusively on the right side of the threshold. For interest, the whole 90% interval is at or above the target.
  • fail: conclusively on the wrong side, at α = 0.05.
  • inconclusive: not enough sends to decide either way. It comes with how many more sends it needs (sends_to_conclude).
  • blocked: interest can't be read because delivery failed.

The range shown next to each rate is a 90% Clopper-Pearson interval. It stays honest at zero replies, where simpler intervals collapse to a single point and claim certainty exactly when there is the least information. Watching it narrow as sends accumulate is the best way to see how much the data can actually tell you.

Use the Signal Check's interest numbers, not the engagement numbers on the campaign dashboard. The dashboard counts read receipts as engagement; the Signal Check counts only replies written by a person.

7.2 Common mistakes

  • A false "doesn't work" from delivery. Reading zero replies as "this segment is dead" while the mail sat in spam. Always read delivery first. Our own July 2026 wave: 337 emails across 13 campaigns got zero replies and zero out-of-office replies (delivery fail, p = 0.034). The cause was a sending IP with no reverse DNS record, not the 13 segments.
  • Peeking at interest. Checking the interest verdict every day and stopping the first time it says pass (or fail). Checked often enough, noise produces a verdict sooner or later. Watch delivery and harm daily; read interest once, at the planned sample.
  • Moving the target. Setting or changing the target after seeing the reply rate turns the test into a description of what happened.
  • Mixing up per-lead and per-send rates. A 10% per-lead reply rate in a two-step sequence shows up as roughly 5–6% per email sent.
  • A contaminated sample. Leads who already heard from you. Keep the saved search's rejected list set to All leads (section 4.1).
  • Changing the email mid-experiment. Any edit to copy, subject or sequence starts a new experiment.
  • Generalising from one email. One template sent through one set of mailboxes is, in effect, one experiment repeated many times. "Doesn't work" means this offer, written this way, to this segment doesn't work, not that the audience can never be reached. The Signal Check shows how many template variants the sends were spread across for this reason.
  • Reading a p-value as the probability that the hypothesis is true. It is not. It is how surprising the data would be if the true rate were exactly the target.
  • Treating the grey zone as a yes. "Not clearly above target" is not "above target".

Appendix A. Procedure for an AI agent

This appendix is for an agent connected to the Sellvance MCP server (/mcp/ws/<workspace>/). Follow the steps in order. Ask the user for approval before any step that creates state or sends email, and show them the parameters first.

A.1 Inputs to collect from the user

Input Required Example
experiment_name yes Logistics ops — order re-keying
Question type yes ICP / offer / feature need
Hypothesis yes segment × persona × offer → reply rate ≥ target
Target reply rate, per email sent yes 0.05
Expected reply rate, per email sent yes 0.10
Sequence shape yes 2 steps
Sending mailboxes yes enabled mailboxes with IMAP connected
Sender name yes a real person's name
Daily enrolment pace (add_limit) yes 20

A.2 Thresholds

Parameter Value
Test Exact binomial per signal, α = 0.05 per tail, 90% Clopper-Pearson interval
Minimum sends before any verdict 30
Planned sample (sends) From the table in section 1 for the target and expected rate
Leads to enrol Planned sample ÷ number of steps
Grey-zone extension At most 50% of the planned sample, once

Rates in API responses are proportions (0.05 = 5%). Workspace threshold settings are percentages (5.0 = 5%).

A.3 Call sequence

Gate: channel fit (before anything else)

  1. Ask the user the six questions from "Gate: is cold outreach the right channel?":

    • Can the ACV not pay for the channel? If the user has their ACV, reply-to-deal conversion and all-in cost per email, compute the break-even threshold (section 2) and treat it as a "yes" when that threshold is above the rate they would act on.
    • Does the buyer not read work email?
    • Is the persona saturated?
    • Is there no trigger?
    • Does the jurisdiction require prior consent?
    • Does the value need a demo to be understood?

    If any answer is yes, stop. Don't create segments, saved searches, sequences or campaigns. Tell the user "Cold outreach is the wrong channel for this ICP" and which condition applies. Continue only when every answer is no.

Preflight

  1. workspaces_mailboxes_list: confirm the chosen mailboxes are enabled and have IMAP connected. Without IMAP, replies and auto-replies are never counted, and every verdict is invalid.
  2. workspaces_mailboxes_health_retrieve for each mailbox, and workspaces_domains_verify_create for each sending domain. Stop and report if any fails.
  3. Target rate: read target_reply_rate on the workspace (a percentage). If it differs from the experiment's target, it must be changed in the workspace settings (the UI, or PATCH /api/v1/workspaces/{sqid}/, which needs account admin rights and the account-level MCP server). If another experiment in this workspace is still running, stop and ask the user: changing the target invalidates it (section 4.2).

Sample

  1. segments_list: find the All leads segment (a single filter with no conditions). If it's missing, create a filter with workspaces_filter_configs_create (conditions = []), then segments_create "All leads" and attach the filter with segments_update.
  2. segments_update on All leads with no field changes. Saving re-evaluates its filter, so every lead already in the workspace becomes a member.
  3. segments_create: <experiment_name> — Prospects and <experiment_name> — Qualified. Build the enrichment chain on Prospects with workspaces_segments_automations_create (the chain from the create-campaign skill); its final qualification step sets positive_move_to_segment to Qualified.
  4. saved_searches_create with add_to_segment = Prospects, rejected_segment = All leads, and is_enabled = false. Autopilot pulls from a saved search whether or not it is enabled; is_enabled only controls the hourly scheduled run. The saved search's own add_limit caps how many companies it adds per day.

Validate by hand

  1. sequences_create and steps_create.
  2. saved_searches_add_now with count = 3–5. Wait for the chain to finish, then check each lead with leads_automation_logs: every output field must be filled and correct.
  3. campaigns_create bound to the Qualified segment, with the chosen mailboxes, is_autopilot_enabled = false, add_limit = 0 and is_unsubscribe_link_enabled = true. With the user's approval, campaigns_activate. With autopilot off and add_limit = 0, nothing is enrolled automatically.
  4. campaigns_add_leads with the qualified validation leads, then campaigns_generate_messages for the first step. Read the generated messages with messages_list / messages_get: every ${...} field must have resolved, and the copy must read as a real person wrote it.
  5. placement_tests_create on one generated message from each sending mailbox, then poll placement_tests_get until the result is no longer pending (up to 4 hours). If any lands in spam or is undelivered, stop and report.
  6. With the user's approval, messages_bulk_approve the validation messages. They count toward the planned sample.

Run

  1. With the user's approval, campaigns_update: is_autopilot_enabled = true and add_limit = the agreed daily pace. Then campaigns_autopilot_run (with the default dry_run = true) to check for blockers such as no_upstream_sources or send_capacity_exhausted.
  2. Every day, campaigns_signals. Read signals.delivery first. If delivery is fail, or signals.harm is fail, campaigns_pause and report the outcome from section 6.
  3. When the number of enrolled leads reaches planned sample ÷ number of steps (campaigns_stats, first-step funnel), campaigns_update with add_limit = 0. Follow-ups continue for enrolled leads.

Verdict

  1. When counts.sends has reached the planned sample and the last step's delay has passed for the last enrolled leads, read signals.interest and apply the stopping rules in section 6. For the grey zone, raise add_limit once to enrol at most 50% more leads, then read again.
  2. Report to the user: the hypothesis, the target, sends, genuine replies, the reply rate with its interval, the delivery and harm statuses, and the outcome in the words of section 6. Say plainly when the result is "doesn't work", "grey zone" or "invalid".

A.4 An agent must not

  • Design or run an experiment when any channel-fit condition applies. Report the reason instead.
  • Report an interest verdict before the planned sample is reached.
  • Interpret interest when delivery is fail or interest is blocked.
  • Change the sequence, steps, mailboxes, segment filters or target rate of a running experiment. Any change ends it.
  • Run an experiment without rejected_segment = All leads on its saved search.
  • Activate a campaign, enable autopilot, or approve messages without the user's explicit approval.
  • Compare two experiments with each other as if they were variants of one test.
  • Describe a p-value as the probability that a hypothesis is true.
  • Present a grey-zone result as "works".

Run your first experiment

Pick one hypothesis, set the target before you send, and let the Signal Check tell you when the answer is in.