---
title: "The GTM Experiment Playbook"
description: "How to find out whether a segment, an offer or a feature idea gets a response from the market, with a few hundred emails: hypothesis, target rate, delivery gate, sample size, stopping rules, and how to read the verdict."
version: 2026-09-15
---

# The GTM Experiment Playbook

This playbook is for B2B teams that don't yet know who to sell to, or what to
sell them. It turns "let's try cold email" into an experiment: one question,
a target decided in advance, a sample size decided in advance, and a yes or
no you can defend.

It is written for two readers at once. A person can read it top to bottom. An
AI agent connected to the Sellvance MCP server can follow it step by step:
every rule is stated as a threshold or a check, and the appendix lists the
tool calls in order.

**The questions an experiment answers:**

- Does this ICP respond, or not?
- Does this offer work, or not?
- Does our ICP need this feature, or not?

**What an experiment gives you:** the answer, including "no", in weeks
rather than months.

**What it does not give you:** leads, meetings, or revenue. Those can happen
along the way. They are not what the experiment is for.

---

## 1. What an answer costs

Each experiment tests one campaign against a **threshold**: the reply rate at
which the answer is "yes, this works". The Signal Check in Sellvance tells you
when the result is conclusively above it, or conclusively below it.

This playbook never tells you what reply rate is normal. Published "average
reply rates" come from the customers of the tools that publish them, and they
mix products, markets and offers that have nothing in common. What it tells
you instead is what it costs to find out:

| What's tested | Threshold | If the true rate is | Sends to a verdict |
|---|---|---|---|
| Delivery (zero out-of-office replies) | 1% | 0% | 299 |
| Interest | 3% | 6% | ~207 |
| Interest | 3% | 8% | ~83 |
| Interest | 3% | 10% | ~42 |
| Interest | 3% | 15% | 30 |
| Interest | 5% | 15% | ~36 |
| Interest | 10% | 25% | ~31 |

The delivery row is exact: after 299 sends with no auto-reply at all, a 1%
baseline is rejected. The interest rows are the sends for an 80% chance of a
"works" verdict, computed from the Signal Check's own test (exact binomial,
90% Clopper-Pearson interval). 30 is the smallest sample the Signal Check will
judge at all.

None of these rows says what reply rate you will get. Each says how many
emails it takes to know. **The stronger your offer is against your threshold,
the cheaper it is to prove.** A precise segment and a message written for it
are what move you down the table.

If the true rate sits right at your target, no sample size will give a clean
verdict. That is a real answer too (section 6, grey zone).

---

## Gate: is cold outreach the right channel?

Cold outreach doesn't work for every ICP. Check these **before** you design an
experiment. They are a hard gate, not advice: if any one of them is true for
your ICP, don't design the experiment. The answer is already "no", and no
amount of copy changes it.

| Condition | Why the channel fails |
|---|---|
| **The ACV can't pay for the channel** | A self-serve product at $20 a month can't pay for domains, mailboxes and the time spent handling replies. Check it with the break-even threshold in section 2: if it's above the rate you would honestly act on, stop. |
| **The buyer doesn't read work email** | Micro-business owners, developers, doctors, tradespeople. Silence tells you about their inbox, not about your offer. |
| **The persona is saturated** | For example a VP of Sales or a CTO at a funded startup: filters, assistants and a learned reflex bury cold email before anyone reads it. |
| **There is no trigger** | If nothing makes the problem urgent at a particular moment, the email always arrives at the wrong time. |
| **The jurisdiction requires prior consent** | Where the law requires opt-in before contact, the sample you can legally reach is zero. |
| **The value can't be understood without a demo** | If the product can't be explained in about 80 words, replies say "didn't get it", not "no thanks". You'd be testing the explanation, not the offer. |

If one of these applies, record which one and stop. Finding that out in an
afternoon is the cheapest answer this playbook can give.

---

## 2. Write the hypothesis

One experiment answers one question. Write it down before you build anything:

> **Segment** × **Persona** × **Offer** → reply rate of at least **Threshold**

- **Segment:** the companies. Narrow enough that the pain is the same across
  all of them. Not "B2B SaaS", but "bootstrapped B2B SaaS, 2–20 people, no
  sales hire".
- **Persona:** the role you email at those companies. Founder, head of sales,
  operations lead.
- **Offer:** the one promise the email makes, and the yes/no action it asks
  for.
- **Threshold:** the reply rate below which you would not act on this. Decide
  it before you see any data, and never move it afterwards.
- **Expected rate:** the rate you want to be able to prove. It sizes the
  experiment (section 5). It is an assumption for sizing, not a forecast.

### Where the threshold comes from

Not from an industry average. The question isn't "what reply rate is normal?"
but "what reply rate makes this channel pay for itself for us?". Cold email
pays off when

> ACV × reply-to-deal conversion × replies > the cost of the sends and your time

so the break-even reply rate per email sent is

> threshold = all-in cost per email ÷ (ACV × reply-to-deal conversion)

A product with a $30,000 ACV where one reply in ten closes breaks even at 1.5%
if each email costs $45 all-in, counting your time. A $600-a-year product with
the same conversion needs 12% even at $7.20 per email. There is no single
right number; there is only yours. The experiment planner calculates it for
you.

What each kind of question holds fixed:

| Question | What the experiment tests | Example hypothesis |
|---|---|---|
| Does this ICP respond? | The segment | Bootstrapped B2B SaaS founders reply to offer X at ≥ 10% |
| Does this offer work? | The promise in the email | Agency owners reply at ≥ 8% to "we find your next five clients" |
| Does our ICP need this feature? | The problem the feature solves, offered before it is built | Ops leads at 50–200 person logistics firms reply at ≥ 10% to "stop re-keying orders from email" |

One experiment changes one thing. If "doesn't work" comes back, change a
single element (the segment, the persona, or the offer) and run a new
experiment on fresh leads. Changing two at once leaves you unable to say which
one mattered.

### Rates are per email sent

The Signal Check counts every email sent, across all steps of the sequence. In
a two-step sequence, everyone who doesn't reply to the first email gets a
second, so the rate per email is roughly half the rate per lead. Set the
target in the same unit the Signal Check uses: **replies per email sent**.

---

## 3. Gate: confirm delivery before you trust any result

A reply rate mixes two things: whether the email **arrived**, and whether it
**persuaded**. Zero replies is equally consistent with a bad offer and with a
spam folder, and the two need opposite fixes. A "doesn't work" verdict is only
valid once delivery is confirmed. This is a hard gate, not a recommendation.

**Before sending:**

1. Every sending domain has SPF, DKIM, DMARC and MX published and verified.
2. Every sending mailbox passes its health check and has IMAP connected.
   Without IMAP, replies are never counted.
3. Run an **inbox placement test** on the real rendered first message, from each
   sending mailbox. Caveat: the test lands in one of your own IMAP mailboxes, so
   it catches authentication failures and spam-filter scores, but "inbox" there
   is not proof of inboxing at Gmail or Outlook.

**While sending:**

4. Read the **delivery** signal on the campaign's Signal Check first.
   Delivery is measured with out-of-office auto-replies: mail servers send them
   without a human reading anything, and Gmail and Outlook suppress them for
   mail classified as spam. So the auto-reply rate reflects inbox placement and
   nothing else.
5. If delivery **fails**, interest is reported as **blocked**. Do not read the
   reply rate. Pause, fix infrastructure, then start again on fresh leads.
6. Zero auto-replies has two possible causes: spam placement, or replies not
   being collected (no IMAP connection). Rule out the second before concluding
   the first.

Auto-reply rates are seasonal: roughly 1–3% in ordinary weeks, 8–12% in
August and late December. The delivery threshold uses your workspace's own
measured rate once it has at least 200 sends of history.

---

## 4. Design

### 4.1 A fresh sample

Every lead in the experiment must be a company nobody from your side has
contacted. A company that already got an email from you is answering a second
email, not your offer.

The database holds 9.5 million companies, so fresh leads are not the
constraint. Keeping them fresh is. Sellvance does it with the saved search's
**rejected list**: a saved search never adds a company that has any lead in
that segment.

1. Check that the workspace has an **All leads** segment: a segment with one
   filter that has no conditions. New workspaces come with one. Create it if
   it's missing.
2. Save that segment once before the experiment starts. Saving re-evaluates
   its filter, so every lead already in the workspace becomes a member. From
   then on, new leads join it as they're added.
3. Set the rejected list of the experiment's saved search to All leads.

Every company that an earlier experiment or campaign touched is now excluded.

### 4.2 One target per workspace

The target reply rate is a workspace setting. Experiments that share a target
can run side by side in one workspace. An experiment with a different target
needs its own workspace, or has to wait until the others finish.

Experiments running side by side are each judged against the target. Don't
compare them with each other: two campaigns that both pass are both "yes",
not a ranking.

### 4.3 Hold everything fixed while it runs

Once the first email of the experiment is sent, don't change:

- the sequence, its steps, subjects or copy
- the sending mailboxes
- the target reply rate
- the segment's filters or enrichment chain

Any of these changes starts a new experiment. The sends before the change
don't count toward it.

---

## 5. Size and run the experiment

1. **Pick the planned sample.** Enter your threshold and the rate you want to
   be able to prove into the experiment planner, or read it off the table in
   section 1. When unsure, size for the lowest rate you would still act on: it
   takes more sends, but the verdict holds either way. Write the number down:
   it's the point where you read the verdict.
2. **Validate by hand first.** Build the campaign with autopilot off, add 3–5
   leads manually, and check everything end to end: the enrichment filled
   every field, the generated emails read well, the placement test lands in
   the inbox. Manual adding exists for exactly this check.
3. **Turn on autopilot.** Once validation passes, enable autopilot and set the
   campaign's daily limit (`add_limit`). From then on, autopilot pulls new
   companies through the saved search, runs the enrichment chain, enrols
   qualified leads, and generates and reviews the emails.
4. **Stop enrolling at the planned sample.** Leads needed ≈ planned sample ÷
   number of steps. When that many leads are enrolled, set `add_limit` to 0.
   Follow-ups to leads already enrolled keep going.
5. **Read the verdict** once the Signal Check shows the planned number of
   sends and the last follow-ups have gone out (section 6).

If the planned sample is more than you can afford, change the experiment, not
the rule: test a narrower segment where you expect a higher reply rate, or a
bolder offer.

---

## 6. Stopping rules, decided in advance

Write these down next to the hypothesis before the first email goes out.

| Outcome | Signal Check at the planned sample | What to do |
|---|---|---|
| **Works** | Delivery passes, interest **passes** | The segment × persona × offer beats the target. Scale it: start a separate campaign with a bigger daily limit and more mailboxes. Nothing on the platform stops you. Keeping it separate leaves the experiment's numbers clean. |
| **Doesn't work** | Delivery passes, interest **fails** | This combination is conclusively below the target. Change one element and run a new experiment on fresh leads. |
| **Grey zone** | Delivery passes, interest still **inconclusive** | The true rate is close to the target. Extend once, by at most 50% of the planned sample. If it is still inconclusive, the answer is "not clearly above target": treat it as no, and move on. |
| **Harmful** | Harm **fails** at any point | Unsubscribes are conclusively above your limit. Pause the campaign. The offer or the list is wrong for these people, whatever the reply rate says. |
| **Invalid** | Delivery **fails**, or interest is **blocked** | No conclusion about the market is possible. Pause, fix delivery, restart on fresh leads. |

Delivery and harm can end an experiment early. Interest is read only at the
planned sample (section 7.2 explains why).

---

## 7. Reading the verdict

### 7.1 The Signal Check

Every campaign, and the workspace as a whole, has a Signal Check with three
signals, each an exact binomial test against a threshold:

| Signal | Counts | Fails when |
|---|---|---|
| **Delivery** | Out-of-office auto-replies | Too few |
| **Interest** | Replies written by a person | Too few |
| **Harm** | Unsubscribes | Too many |

Each signal has one of four statuses:

- **pass**: conclusively on the right side of the threshold. For interest, the
  whole 90% interval is at or above the target.
- **fail**: conclusively on the wrong side, at α = 0.05.
- **inconclusive**: not enough sends to decide either way. It comes with how
  many more sends it needs (`sends_to_conclude`).
- **blocked**: interest can't be read because delivery failed.

The range shown next to each rate is a 90% Clopper-Pearson interval. It stays
honest at zero replies, where simpler intervals collapse to a single point and
claim certainty exactly when there is the least information. Watching it
narrow as sends accumulate is the best way to see how much the data can
actually tell you.

Use the Signal Check's interest numbers, not the engagement numbers on the
campaign dashboard. The dashboard counts read receipts as engagement; the
Signal Check counts only replies written by a person.

### 7.2 Common mistakes

- **A false "doesn't work" from delivery.** Reading zero replies as "this
  segment is dead" while the mail sat in spam. Always read delivery first.
  Our own July 2026 wave: 337 emails across 13 campaigns got zero replies and
  zero out-of-office replies (delivery fail, p = 0.034). The cause was a
  sending IP with no reverse DNS record, not the 13 segments.
- **Peeking at interest.** Checking the interest verdict every day and stopping
  the first time it says pass (or fail). Checked often enough, noise produces a
  verdict sooner or later. Watch delivery and harm daily; read interest once,
  at the planned sample.
- **Moving the target.** Setting or changing the target after seeing the reply
  rate turns the test into a description of what happened.
- **Mixing up per-lead and per-send rates.** A 10% per-lead reply rate in a
  two-step sequence shows up as roughly 5–6% per email sent.
- **A contaminated sample.** Leads who already heard from you. Keep the saved
  search's rejected list set to All leads (section 4.1).
- **Changing the email mid-experiment.** Any edit to copy, subject or sequence
  starts a new experiment.
- **Generalising from one email.** One template sent through one set of
  mailboxes is, in effect, one experiment repeated many times. "Doesn't work"
  means *this offer, written this way, to this segment* doesn't work, not that
  the audience can never be reached. The Signal Check shows how many template
  variants the sends were spread across for this reason.
- **Reading a p-value as the probability that the hypothesis is true.** It is
  not. It is how surprising the data would be if the true rate were exactly the
  target.
- **Treating the grey zone as a yes.** "Not clearly above target" is not
  "above target".

---

## Appendix A. Procedure for an AI agent

This appendix is for an agent connected to the Sellvance MCP server
(`/mcp/ws/<workspace>/`). Follow the steps in order. Ask the user for approval
before any step that creates state or sends email, and show them the
parameters first.

### A.1 Inputs to collect from the user

| Input | Required | Example |
|---|---|---|
| `experiment_name` | yes | `Logistics ops — order re-keying` |
| Question type | yes | ICP / offer / feature need |
| Hypothesis | yes | segment × persona × offer → reply rate ≥ target |
| Target reply rate, **per email sent** | yes | `0.05` |
| Expected reply rate, per email sent | yes | `0.10` |
| Sequence shape | yes | 2 steps |
| Sending mailboxes | yes | enabled mailboxes with IMAP connected |
| Sender name | yes | a real person's name |
| Daily enrolment pace (`add_limit`) | yes | `20` |

### A.2 Thresholds

| Parameter | Value |
|---|---|
| Test | Exact binomial per signal, α = 0.05 per tail, 90% Clopper-Pearson interval |
| Minimum sends before any verdict | 30 |
| Planned sample (sends) | From the table in section 1 for the target and expected rate |
| Leads to enrol | Planned sample ÷ number of steps |
| Grey-zone extension | At most 50% of the planned sample, once |

Rates in API responses are proportions (`0.05` = 5%). Workspace threshold
settings are percentages (`5.0` = 5%).

### A.3 Call sequence

**Gate: channel fit (before anything else)**

0. Ask the user the six questions from "Gate: is cold outreach the right
   channel?":
   - Can the ACV not pay for the channel? If the user has their ACV,
     reply-to-deal conversion and all-in cost per email, compute the
     break-even threshold (section 2) and treat it as a "yes" when that
     threshold is above the rate they would act on.
   - Does the buyer not read work email?
   - Is the persona saturated?
   - Is there no trigger?
   - Does the jurisdiction require prior consent?
   - Does the value need a demo to be understood?

   If any answer is yes, **stop.** Don't create segments, saved searches,
   sequences or campaigns. Tell the user "Cold outreach is the wrong channel
   for this ICP" and which condition applies. Continue only when every answer
   is no.

**Preflight**

1. `workspaces_mailboxes_list`: confirm the chosen mailboxes are enabled and
   have IMAP connected. Without IMAP, replies and auto-replies are never
   counted, and every verdict is invalid.
2. `workspaces_mailboxes_health_retrieve` for each mailbox, and
   `workspaces_domains_verify_create` for each sending domain. Stop and report
   if any fails.
3. Target rate: read `target_reply_rate` on the workspace (a percentage). If it
   differs from the experiment's target, it must be changed in the workspace
   settings (the UI, or `PATCH /api/v1/workspaces/{sqid}/`, which needs account
   admin rights and the account-level MCP server). If another experiment in
   this workspace is still running, stop and ask the user: changing the target
   invalidates it (section 4.2).

**Sample**

4. `segments_list`: find the **All leads** segment (a single filter with no
   conditions). If it's missing, create a filter with
   `workspaces_filter_configs_create` (`conditions` = `[]`), then
   `segments_create` "All leads" and attach the filter with `segments_update`.
5. `segments_update` on All leads with no field changes. Saving re-evaluates
   its filter, so every lead already in the workspace becomes a member.
6. `segments_create`: `<experiment_name> — Prospects` and
   `<experiment_name> — Qualified`. Build the enrichment chain on Prospects
   with `workspaces_segments_automations_create` (the chain from the
   `create-campaign` skill); its final qualification step sets
   `positive_move_to_segment` to Qualified.
7. `saved_searches_create` with `add_to_segment` = Prospects,
   `rejected_segment` = All leads, and `is_enabled` = false. Autopilot pulls
   from a saved search whether or not it is enabled; `is_enabled` only controls
   the hourly scheduled run. The saved search's own `add_limit` caps how many
   companies it adds per day.

**Validate by hand**

8. `sequences_create` and `steps_create`.
9. `saved_searches_add_now` with `count` = 3–5. Wait for the chain to finish,
   then check each lead with `leads_automation_logs`: every output field must
   be filled and correct.
10. `campaigns_create` bound to the **Qualified** segment, with the chosen
    `mailboxes`, `is_autopilot_enabled` = false, `add_limit` = 0 and
    `is_unsubscribe_link_enabled` = true. With the user's approval,
    `campaigns_activate`. With autopilot off and `add_limit` = 0, nothing is
    enrolled automatically.
11. `campaigns_add_leads` with the qualified validation leads, then
    `campaigns_generate_messages` for the first step. Read the generated
    messages with `messages_list` / `messages_get`: every `${...}` field must
    have resolved, and the copy must read as a real person wrote it.
12. `placement_tests_create` on one generated message from each sending
    mailbox, then poll `placement_tests_get` until the result is no longer
    `pending` (up to 4 hours). If any lands in spam or is undelivered, stop and
    report.
13. With the user's approval, `messages_bulk_approve` the validation messages.
    They count toward the planned sample.

**Run**

14. With the user's approval, `campaigns_update`: `is_autopilot_enabled` =
    true and `add_limit` = the agreed daily pace. Then
    `campaigns_autopilot_run` (with the default `dry_run` = true) to check for
    blockers such as `no_upstream_sources` or `send_capacity_exhausted`.
15. Every day, `campaigns_signals`. Read `signals.delivery` first. If delivery
    is `fail`, or `signals.harm` is `fail`, `campaigns_pause` and report the
    outcome from section 6.
16. When the number of enrolled leads reaches planned sample ÷ number of steps
    (`campaigns_stats`, first-step funnel), `campaigns_update` with
    `add_limit` = 0. Follow-ups continue for enrolled leads.

**Verdict**

17. When `counts.sends` has reached the planned sample and the last step's
    delay has passed for the last enrolled leads, read `signals.interest` and
    apply the stopping rules in section 6. For the grey zone, raise `add_limit`
    once to enrol at most 50% more leads, then read again.
18. Report to the user: the hypothesis, the target, sends, genuine replies, the
    reply rate with its interval, the delivery and harm statuses, and the
    outcome in the words of section 6. Say plainly when the result is "doesn't
    work", "grey zone" or "invalid".

### A.4 An agent must not

- Design or run an experiment when any channel-fit condition applies. Report
  the reason instead.
- Report an interest verdict before the planned sample is reached.
- Interpret `interest` when `delivery` is `fail` or `interest` is `blocked`.
- Change the sequence, steps, mailboxes, segment filters or target rate of a
  running experiment. Any change ends it.
- Run an experiment without `rejected_segment` = All leads on its saved search.
- Activate a campaign, enable autopilot, or approve messages without the user's
  explicit approval.
- Compare two experiments with each other as if they were variants of one test.
- Describe a p-value as the probability that a hypothesis is true.
- Present a grey-zone result as "works".
