| Takeaway | Detail |
|---|---|
| Require any RL cadence tool to show its suppression rate before adoption. | Suppression rate is the percentage of scheduled sends the tool chose NOT to send; the reader rule treats below 15% as a red flag. |
| Treat the 18% reply-rate lift as a suppression result, not a send-more result. | The thesis ties the lift to the prior capped-vs-static-vs-blast analysis and to learning when to suppress sends. |
| Benchmark the reward function against systems with known optimal policies. | Use stochastic-converse-optimality benchmarking, arXiv:2603.17631, to generate systems where the optimal policy is known before trusting the reward. |
| Validate safe exploration and constraints in standardized realistic or cost-efficient testbeds. | Use RL2Grid with Grid2Op for a standardized realistic power-grid benchmark, the (N,K)-Puzzle, arXiv:2403.07191, as a cost-efficient RL testbed, and benchmark safe exploration in deep RL. |
This guide shows how to evaluate reinforcement-learning send-time tools by their suppression behavior, not send volume. It translates the 18% reply-rate lift into checks for reward design, benchmarking, and safe constrained deployment.

How RL cadence actually suppresses sends
The mechanism that produces suppression is the reward function, and it is the one thing most cadence vendors will not show you.
In practical real-world RL work — the constrained-optimization line of research exemplified by the Optlayer paper catalogued by Springer, alongside the safe-exploration benchmarking tradition that runs back to Riedmiller's 2005 work — agents are typically not rewarded for the target outcome alone; constraints and penalties shape the objective. For email cadence, that means the reward you should ask to see is a reply term net of a contact-fatigue penalty, and the penalty is what makes holding a send a rational action rather than a missed opportunity. Whether a given vendor's reward actually contains such a term is exactly what you must verify, not assume.
If a tool describes its objective as "maximize engagement," you are looking at an unconstrained reward, and an unconstrained reward has no reason to ever suppress.Ask which state inputs the policy actually reads. The three that carry the suppression signal are recency of last touch, prior open-without-reply streaks, and time-of-day engagement history for that specific contact. Calendar heuristics — day-of-week rules, fixed gaps between sends — are not state; they are constants, and a policy conditioned on constants cannot learn that the same contact should be mailed on Tuesday and held on Thursday. The policy's job is a per-contact, per-window send-or-hold decision, which means the suppression rate is an emergent property of the state representation, not a setting someone dials in.
Here is the check that separates a real policy from a scheduler wearing RL branding. Request the suppression rate: the percentage of scheduled sends the system chose not to send. Then request the reward decomposition — the reply term and the fatigue penalty term, separately, with the penalty's weight. A vendor who can produce the first number but not the second is reporting an outcome it cannot explain, and an outcome you cannot explain is one you cannot audit when reply rates drift.
| What to request | What a genuine suppression policy shows | What a disguised scheduler shows |
|---|---|---|
| Suppression rate | A measured share of scheduled sends withheld | A single figure, or none |
| Reward decomposition | Reply term plus a named fatigue penalty | "Engagement score," undefined |
| State inputs | Recency, open-without-reply streaks, time-of-day history | Calendar cadence rules |
| Decision granularity | Per contact, per window | Per campaign |
Treat a suppression rate below 15% as evidence the system is a scheduler, not a policy. The reasoning is arithmetic rather than tribal: if the agent withholds almost nothing, its send decisions are effectively fixed in advance, and a fixed sequence of sends is exactly what the static and blast arms of the earlier comparison already were. A tool that suppresses nothing cannot be the source of a lift that the static arm did not produce.
One caution on benchmarking. The arXiv work on stochastic converse optimality (2603.17631) makes the point that RL comparisons are critically sensitive to environmental design, reward structure, and stochasticity in both learning and dynamics — which is why a vendor's own benchmark, run on its own environment, is weak evidence. Ask for the reward structure and the suppression rate on your data, in your sending environment, before you trust any lift number at all.

The evidence: where 18% holds and why
The 18% reply-rate lift that anchors this argument comes from a single measurement: our prior capped-versus-static-versus-blast analysis referenced at the top of this piece. That provenance matters more than the number. It is one team's measured result on one corpus of sequences, not an industry constant, and no vendor should be allowed to quote it as though it were a published benchmark. When a tool claims an 18% lift, ask which sequences it was measured against, over what window, and whether the comparison held suppression constant. If the vendor cannot answer, the figure is decoration.
The method point that should govern how you read any such claim is established in the arXiv paper on benchmarking reinforcement learning via stochastic converse optimality (arXiv:2603.17631), which shows that outcomes and benchmarking of different RL approaches are critically sensitive to environmental design, reward structures, and the stochasticity inherent in both algorithmic learning and environmental dynamics.
The practical consequence for a buyer is direct: a vendor's lift number is not a property of the algorithm, it is a property of the algorithm plus the environment it was tuned in. Two tools running nominally the same method can diverge sharply once the reward structure or the noise profile changes, which is exactly what happens when you move from a vendor's demo corpus to your own list.That is why the 18% figure should be treated as a hypothesis to replicate, not a specification to purchase. The replication check is the suppression rate: the percentage of scheduled sends the system chose not to send. A cadence that earns its lift by suppressing sends will show a suppression rate that moves with list conditions; a cadence that earns its lift by volume will show a suppression rate near zero. Treat a suppression rate below 15% as evidence the tool is a send-more engine wearing RL vocabulary.
The convergence across independent RL-benchmarking work supports the same reading. The benchmarking literature catalogued across sources including RL2Grid's standardized power-grid benchmark, the (N, K)-Puzzle cost-efficient testbed for evaluating RL algorithms at scale, and demand-response and portfolio-optimization benchmarking studies all share one structural feature: they fix the environment and the reward before comparing policies, precisely because unfixed comparisons do not transfer. None of these sources reports a reply-rate lift for sales email, and none should be cited as if it did. What they establish is the discipline your vendor evaluation needs.
| Source | What it benchmarks | Transferable check |
|---|---|---|
| mm-ais.com / Dawson, Sept 2026 | Capped vs static vs blast sequences | Ask for the suppression rate behind the 18% |
| arXiv:2603.17631 (Ibrahim et al., 2026) | RL benchmarking under stochastic converse optimality | Reward structure and stochasticity drive outcomes |
| RL2Grid (alphaXiv) | RL in power grid operations | Fixed environment before policy comparison |
| (N, K)-Puzzle, arXiv:2403.07191 | Cost-efficient RL testbed at scale | Standardized testbed before performance claims |
Read the table as a single instruction: before adopting any RL cadence tool, require the suppression rate and require the reward structure. The 18% holds only where those two are disclosed and reproducible on your list. Where they are not, the number is a marketing artifact, and the burden of proof stays with the vendor.

Static vs capped vs RL cadence compared
Put the three cadences side by side and the comparison stops being about reply rate alone. A static schedule sends on a fixed day-and-time cadence with zero suppression: every scheduled touch goes out, nothing is held back. That makes it the easiest cadence to audit, because the send log and the plan are identical, and it wins on predictability of volume — you can forecast inbox load per contact without knowing anything about the contact. It loses on reply rate, and it loses for a structural reason: it has no mechanism to withhold a touch that the contact's history says will annoy rather than persuade.
The capped heuristic is the control that matters, and it is the one most teams skip. It combines a hard cap on touches per contact per period with intent-gated sends, so a contact who has shown no engagement signal is not queued at all. In our prior capped-versus-static-versus-blast analysis, this heuristic already captured most of the lift — which means the honest benchmark for any RL policy is the capped heuristic, not the static schedule. If a vendor compares its RL cadence to a static baseline and reports a large gain, ask for the capped comparison instead. Beating static is table stakes; beating capped is the claim.
The RL policy learns per-contact send-or-hold decisions rather than applying one rule to everyone. It wins when contact history is rich — roughly 30 or more prior touches per segment — because that is where per-contact signal exists to learn from. It loses when it overfits sparse data, which is the common failure mode: a policy trained on thin history will confidently suppress sends it should have made, or make sends it should have suppressed, and the reply-rate gain evaporates. The tie-break condition is data density, not algorithm choice.
| Cadence | Suppression | Reply rate | Volume predictability | Auditability |
|---|---|---|---|---|
| Static schedule | None | Lowest | Highest | Easiest |
| Capped heuristic | Cap plus intent gate | Most of the lift | High | Moderate |
| RL policy | Learned per contact | Highest when history is rich | Lowest | Hardest |
Winner: the RL policy, but only under the tie-break condition — 30 or more prior touches per segment. Below that density, the capped heuristic is the better deployment, because it delivers most of the gain with a rule you can read and defend. Before adopting any RL cadence tool, require it to show you its suppression rate: the percentage of scheduled sends it chose not to send. Treat a suppression rate below 15% as evidence it is a static schedule wearing a new label, since a policy that suppresses almost nothing is not making per-contact decisions.
The benchmark framework in arXiv:2603.17631 makes the same point formally: RL outcomes are critically sensitive to environmental design and reward structure, so a policy's headline number means little without the suppression behavior behind it.

Costs and the numbers that matter
Every deployment of an RL cadence policy carries three costs, and you should price all three before you sign anything: the data cost of training, the exploration cost of early sends it holds back, and the arithmetic gap between your baseline and the promised lift. Vendors quote the lift; almost none quote the other two. This section is the only place in this piece where those costs get numbers attached.
Start with exploration. An RL policy learns by sometimes withholding sends on contacts who would have replied — that is how it discovers which holds pay off.
The practical consequence is a measurable dip in reply rate during the early weeks, and the benchmarking literature is blunt about why: outcomes are critically sensitive to environmental design, reward structures, and stochasticity in both the algorithm and the environment, which is exactly the framing of the arXiv framework for benchmarking RL via stochastic converse optimality (arXiv:2603.17631).
Translation for your pipeline: judge the policy only after it has stopped exploring. Set your evaluation window at a minimum of two full cadence cycles per segment before you compare against your static baseline, and ask the vendor in writing what early-weeks dip their existing customers saw.Then run the arithmetic yourself, because the lift is multiplicative, not additive. If your static baseline is 40 replies per 1,000 sends, an 18% lift means roughly 47 replies per 1,000 — 40 × 1.18 = 47.2. On 20,000 sends per month, that is 800 baseline replies versus about 940, or 140 extra replies per month. If your baseline is 25 per 1,000, the same 18% yields only 29.5 per 1,000 — the absolute gain shrinks with your baseline, so compute your own number first, not the vendor's example.
| Cost item | What to quantify | Check to run |
|---|---|---|
| Exploration dip | Reply-rate decline while the policy learns | Require the vendor's observed early-weeks dip; evaluate over two full cadence cycles per segment |
| Suppression rate | Percentage of scheduled sends the policy chose not to send | Ask for the number directly; a low suppression rate suggests the policy is not actually learning to hold sends |
| Lift arithmetic | Extra replies per month at your volume | Multiply your baseline replies per 1,000 by 1.18, then by your monthly send volume |
The suppression rate deserves its own line in the contract because it is the cheapest audit available. A policy that suppresses almost nothing is not trading short-term sends for long-term learning — it is a static cadence with a machine-learning label and a higher invoice. The benchmarking work above exists precisely because RL performance claims are sensitive to how the evaluation is set up; your defense is to fix the evaluation setup before the vendor does.
Finally, budget the data. A policy learning per-segment send timing needs enough reply events per segment to distinguish signal from noise, and segments with thin reply history will converge slowly or not at all. Ask how many reply events per segment the vendor considers a floor, and exclude segments below it from the initial rollout rather than letting the policy learn on contacts too quiet to teach it anything.

What the 18% does not establish
The 18% reply-rate lift is a real measurement, but it is one team's result under one reward structure. Per arXiv:2603.17631, RL outcomes and benchmarking performance are "critically sensitive to environmental design, reward structures, and stochasticity inherent in both algorithmic learning and environmental dynamics."
That is the paper's central warning, and it applies directly to cadence: change the fatigue penalty, change the list, change the send volume, and the learned policy changes with it. The number does not transfer to your list automatically, and any vendor who presents it as a portable constant is skipping the part where you verify it on your own data.The cleanest way to see the limit is to find where the rule breaks. Two edge cases matter most, and they are the ones a pilot will surface first.
| Edge case | Why the RL cadence rule breaks | What still wins |
|---|---|---|
| Cold outbound, no history | No recency or engagement state exists, so the policy has nothing to learn from; suppression decisions have no signal behind them | The capped heuristic |
| Small lists (under roughly 5,000 contacts) | Per-contact state estimates are too noisy to support a stable suppression decision | The capped heuristic |
| Mature list with engagement history | Suppression has enough signal to act on | The RL cadence rule |
Cold outbound is the first break. With no recency or engagement state, the policy has nothing to learn from — there is no per-contact history to suppress against, so the learned rule degenerates toward whatever prior the reward function encodes. The capped heuristic still wins here because it does not depend on state that does not exist yet.
Small lists are the second break. Under roughly 5,000 contacts, per-contact state estimates are too noisy for the policy to distinguish a genuine suppression signal from variance. The rule does not fail loudly; it fails quietly, suppressing sends for reasons that will not replicate. The capped heuristic wins again, because it does not require per-contact estimation at all.
The practical rule follows from the reader rule this piece serves: before adopting any RL cadence tool, require it to show you its suppression rate — the percentage of scheduled sends it chose not to send — and treat a suppression rate below 15% as evidence it is a send-more tool wearing an RL label. Then run the edge-case check above on your own list. If you are in cold outbound or under the small-list line, the capped heuristic is the correct default, and the 18% is not your number.

Auditing a vendor's RL claim
Auditing a vendor's RL claim starts with three inputs you assemble before the first demo call: your last 90 days of sends, replies, and unsubscribes broken out by segment; a capped-heuristic baseline replayed on that same historical data; and the vendor's suppression-rate export — the percentage of scheduled sends the system chose not to send. If the vendor cannot produce that export, the audit ends there. A cadence engine that cannot show you what it held back is asking you to trust a black box, and the suppression rate is the one number that separates genuine hold decisions from a repackaged send scheduler.
Checkpoint 1 is the replay. Run your capped heuristic over the historical window and record replies per 1,000 sends for each segment. That replayed figure — not your old static schedule, and not the vendor's own before-and-after slide — is the number the vendor must beat. The distinction matters because a static-to-RL comparison flatters any adaptive system; a capped-to-RL comparison isolates the value of the learned suppression logic. Ask the vendor to run the same replay on the same window and hand you the per-segment output. If their methodology cannot reproduce a capped baseline, you have no common yardstick, and the claimed lift is unverifiable.
Checkpoint 2 is the suppression rate itself. Below 15% suppression, the system is not making meaningful hold decisions — flag it. A cadence that suppresses fewer than roughly one in seven scheduled sends is behaving like a throttle with a floor, not a policy learning when silence outperforms contact. This threshold is a screening rule, not a verdict: a low rate on a small, high-intent segment may be defensible, but a low rate across every segment is evidence the "RL" label is cosmetic.
Two supporting checks close the audit. First, confirm the reward function is disclosed in mechanism terms — reply-rate reward net of a fatigue penalty — since the benchmarking literature (Ibrahim et al., arXiv:2603.17631, 2026) shows RL outcomes are critically sensitive to reward structure, meaning an undisclosed reward makes the lift unattributable. Second, verify the suppression export reconciles with your own send logs: the count of scheduled sends minus actual sends should equal the vendor's reported suppressions, segment by segment.
| Audit checkpoint | What you request | Pass condition |
|---|---|---|
| Inputs | 90-day sends/replies/unsubscribes by segment; capped baseline replay; suppression-rate export | All three delivered before contract |
| Checkpoint 1 | Replayed capped heuristic, replies per 1,000 sends | Vendor beats the replayed baseline, not your static schedule |
| Checkpoint 2 | Suppression rate by segment | At or above 15%; below 15% flagged |
| Reconciliation | Scheduled minus actual sends vs. reported suppressions | Counts match segment by segment |
Run this sequence before any pilot converts to a contract. The vendor that passes all four rows has shown you its hold logic; the one that stalls at row one has told you something too.
Worked Example: Run the Numbers
Take one concrete case. A five-touch sequence is scheduled to go to 4,000 contacts over ten business days, two touches per week, with the first send on a Monday. That is 20,000 scheduled sends in total (5 × 4,000). This is an illustration, not a benchmark result — the arithmetic is what matters, and you can swap in your own list size and touch count.
Now apply a suppression rate. If the cadence tool suppresses 20% of scheduled sends, it holds back 4,000 of the 20,000 and actually delivers 16,000. At a 2% reply rate on delivered mail, that yields 320 replies (16,000 × 0.02). A static cadence that delivers all 20,000 at the same 2% reply rate yields 400 replies (20,000 × 0.02). On raw reply count alone, suppression looks like a loss of 80 replies. That comparison is the trap, because it assumes the reply rate is fixed.
The lift only appears when suppression changes the reply rate. Suppose the suppressed sends were the ones landing on already-fatigued contacts, and removing them raises the reply rate on delivered mail from 2% to 2.4%. Now 16,000 delivered × 0.024 = 384 replies. Still short of 400. Push the delivered reply rate to 2.5% and you get 400 replies (16,000 × 0.025) — exactly break-even with the static send. Above 2.5%, suppression wins on replies while sending 20% less mail. That 2.5% is this example's break-even trigger: the delivered reply rate at which a 20% suppression rate matches an unsuppressed 2% baseline.
Run the same arithmetic at the reader-rule floor. A tool suppressing only 10% holds back 2,000 sends and delivers 18,000. To match 400 replies, delivered mail must hit 2.22% (400 ÷ 18,000 = 0.0222). To beat it, higher. So a low suppression rate forces the tool to clear a higher reply-rate bar to justify itself — which is exactly why a suppression rate below 15% should be treated as a signal the tool is not suppressing enough to matter.
Two checks before you accept any vendor's number. First, ask for the suppression rate as a percentage of scheduled sends, not as a count — counts hide list size. Second, ask for the delivered reply rate, not the blended rate across suppressed and delivered mail; blending flatters the result. The benchmarking literature is explicit that RL outcomes are critically sensitive to reward structure and environmental stochasticity (Ibrahim et al., arXiv:2603.17631), so a single headline lift without the suppression rate and the delivered reply rate is not reproducible. Recompute the break-even yourself: baseline reply rate ÷ (1 − suppression rate). At 2% baseline and 20% suppression, that is 0.02 ÷ 0.8 = 2.5% — the same trigger, derived independently.
Decision rules for adopting RL cadence
Adoption decisions should be made on suppression evidence, not on lift claims. The first gate is list size. If your active contact list is under roughly 5,000 reachable addresses, skip reinforcement learning entirely and implement a capped heuristic instead: a fixed ceiling on sends per contact per cycle, enforced in your sequencing tool. Below that density, the policy has too few reply events per send-time bucket to learn a stable suppression boundary, and a cap will match or beat it. Ask the vendor for the number of positive reply events per decision point per week; if that number is small, the policy is guessing.
The second gate is disclosure. If a vendor cannot show you its reward function and its suppression rate, treat the lift claim as unverifiable and walk. The 2026 benchmarking literature is explicit that reward structure determines the result: Ibrahim and colleagues, in arXiv:2603.17631, show that outcomes and performance comparisons are critically sensitive to environmental design, reward structures, and stochasticity, which is why they build a framework for systems with known optimal policies. A vendor that hides the reward function is asking you to accept a number you cannot reproduce. Require the reward definition in writing, including how fatigue is penalized, and require the suppression rate as a standing metric.
Suppression rate is the number that matters at this gate: the percentage of scheduled sends the policy chose not to send. Treat a suppression rate below 15% as evidence the tool is a scheduler with a new label, not a cadence policy. A system that suppresses almost nothing is not learning when to stay quiet; it is learning when to fire. Ask for the rate per segment, not blended across the account, because a high-volume segment can mask near-zero suppression elsewhere.
The third gate is a shadow-mode test. Run the policy in shadow mode against your capped baseline for two full cycles, with the policy's send decisions logged but not executed. Compare replies per 1,000 sends, not total replies, so volume cannot flatter the result. If the policy beats the capped baseline on that metric over both cycles, roll out to that segment only. If it ties or loses, keep the cap and re-test after your list or reply volume changes.
Finally, apply the ethics gate from automated-persuasion research before any rollout. Suppression is a persuasion decision made on a recipient's behalf, so require that the policy's objective is the recipient's reply, not merely the sender's volume, and that opt-out and frequency preferences are hard constraints rather than penalties the policy can trade away. If the vendor cannot describe how recipient preferences constrain the reward, do not deploy.
| Gate | Condition | Action |
|---|---|---|
| List density | Active list under ~5,000 | Use a capped heuristic; skip RL |
| Disclosure | No reward function or suppression rate shown | Treat lift as unverifiable; walk |
| Suppression | Suppression rate below 15% | Reject as a scheduler, not a policy |
| Shadow test | Beats capped baseline on replies per 1,000 sends over two cycles | Roll out to that segment only |
| Ethics | Recipient preferences not hard constraints | Do not deploy |
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Before signing any RL cadence contract, ask the vendor for its suppression rate — the share of scheduled sends the tool chose NOT to send — and put the number in writing. | This is the canonical decision rule: without the suppression rate you cannot tell an RL cadence from a static scheduler wearing an RL label. |
| 2 | If the disclosed suppression rate falls below the 15% threshold, walk away or demand a live audit of the policy's send/suppress decisions. | A sub-15% suppression rate is the red flag defined in the rule — it signals the tool is sending on a fixed schedule, not learning when to hold back. |
| 3 | Re-read the capped-vs-static-vs-blast analysis above and reclassify the 18% reply-rate lift as a suppression result, not a send-more result. | The thesis ties the lift to learning when to suppress sends; misreading it as volume-driven will send you shopping for the wrong tool. |
| 4 | Benchmark the vendor's reward function against systems with known optimal policies using stochastic-converse-optimality benchmarking, arXiv:2603.17631. | You cannot trust a reward signal you have never validated against a system where the optimal policy is known in advance. |
| 5 | Validate safe exploration and constraints in standardized realistic or cost-efficient testbeds — for power-grid-style constraint testing, use RL2Grid with Grid2Op. | Suppression decisions that ignore constraints will not survive production; standardized testbeds expose that before you deploy. |
| 6 | Re-run the suppression-rate check at the 90-day mark and compare the observed rate against the vendor's original disclosure. | A disclosed rate that drifts downward after onboarding is the same static-scheduler failure mode reappearing once the sales pressure is off. |
Frequently Asked Questions
What suppression rate should make me reject an RL cadence tool?
The reader rule treats a suppression rate below 15% as a red flag, so require any RL cadence tool to show its suppression rate before adoption.
Is the 18% reply-rate lift a sign I should send more email?
No — treat the 18% reply-rate lift as a suppression result, not a send-more result.
How can I trust a reward function when I don't know the optimal policy?
Use stochastic-converse-optimality benchmarking, arXiv:2603.17631, to generate systems where the optimal policy is known before trusting the reward.
What standardized realistic testbed can I use to validate safe exploration in RL cadence tools?
Use RL2Grid with Grid2Op for a standardized realistic power-grid benchmark.
What cost-efficient testbed exists for benchmarking RL cadence systems?
The (N,K)-Puzzle, arXiv:2403.07191, serves as a cost-efficient RL testbed.
What actually produces the send suppression in an RL cadence tool?
The mechanism that produces suppression is the reward function, and it is the one thing most cadence vendors will not show you.
Quick answers
| What should you require any RL cadence tool to show before adoption? | Require any RL cadence tool to show its suppression rate before adoption. |
| What is suppression rate? | Suppression rate is the percentage of scheduled sends the tool chose NOT to send. |
| What does the reader rule treat as a red flag? | The reader rule treats below 15% as a red flag. |
| How should the 18% reply-rate lift be treated? | Treat the 18% reply-rate lift as a suppression result, not a send-more result. |
| What should be used to generate systems where the optimal policy is known before trusting the reward? | Use stochastic-converse-optimality benchmarking, arXiv:2603.17631, to generate systems where the optimal policy is known before trusting the reward. |
Also worth reading: Business email sequences 2026: Reinforcement Learning (RL) reply-rate lift vs 4-send cap: Business email sequences 2026: Reinforcement · Email follow up sequence 2026: Reinforcement Learning (RL) vs 5-touch send or stop: Email follow up sequence 2026: · How to get a free personal email domain for your custom address: How to get a free