Skip to content
iMessage APIs
Measurement7 min read

A/B testing business text messages without fooling yourself

At 200 sends a month you will never detect a 3% lift, and chasing one will make your messaging worse. Test the things that move numbers by a third.

Most A/B testing advice assumes ecommerce volume. A clinic sending 400 reminders a month cannot detect a small difference in any reasonable timeframe, and the usual response — declaring a winner anyway — is worse than not testing, because it produces confident false beliefs that compound.

What you can and cannot detect

Monthly sendsRealistically detectableSo test
Under 200Very large differences onlyWhether to send at all; timing by hours
200–1,000Roughly a third or moreAsk structure, message length, send window
1,000–5,000Around 10–15%Copy variants, CTA wording
Over 5,000Smaller effectsClassical A/B on most things

Note the left column is sends per *month*, and you need several weeks for a clean read. If you are in the first row, stop testing copy and start testing whether a play exists at all — that difference is enormous and you can see it.

Test big things

  • Play versus no play. Half the list gets the reminder, half does not. This is the highest-value test any small business can run and almost nobody runs it.
  • Timing in hours, not minutes. Evening-before versus morning-of is a large effect. 5:15pm versus 5:45pm is not.
  • Closed ask versus open ask. 'Tuesday or Thursday?' against 'let us know when suits' routinely moves reply rate by a lot.
  • Named sender versus business name. Often larger than any copy change.
  • One message versus two. Frequency effects are big and usually negative.

Set it up so the result is readable

Use utm_content for the variant and keep everything else identical. If the campaign name changes between variants you have made the two arms incomparable in your reporting.

text
Variant A  ?utm_source=imessage&utm_medium=messaging
           &utm_campaign=appointment-reminder&utm_content=closed-ask

Variant B  ?utm_source=imessage&utm_medium=messaging
           &utm_campaign=appointment-reminder&utm_content=open-ask

Same campaign. Same source and medium. One thing different.

Split by contact, not by day. Splitting Tuesday against Wednesday measures the day of the week, not your copy. Conventions for the rest of the tagging are in UTM tracking for messaging.

Measure the outcome, not the proxy

Read rate is not the result

It is tempting because it moves fast and always looks like progress. But a message can be read more and still book fewer appointments. Pick the outcome first — confirmations, bookings, recovered carts — and use read rate only to explain *why* the outcome moved. That diagnostic split is covered in read receipts.

Two honest rules

  1. Decide the sample and the metric before you start. Stopping when the result looks good is how noise becomes strategy.
  2. If the difference is not obvious, there is no difference. At small-business volume, a result you have to squint at is a result you should discard. Keep the simpler variant and go test something bigger.
testinganalyticscopywriting