Whether segmented delivery is working can be estimated by one act of counting. Set the number of segments you currently have beside the number of distinct kinds of content you actually produce.
Where the two match, the classification is functioning as a unit of production. With six segments and two kinds of content, four of the segments are receiving the same thing. Segments receiving the same thing carry no meaning as divisions.
This count bites because the two often do not match. And the mismatch does not appear in the response figures. If the response does not rise, the reading “the division is too coarse” emerges and the segments multiply. The added segments also receive the same content, so the next reading is that it is coarser still.
The reason for the mismatch is order. The division gets settled first and the content is produced afterwards. Content grows more slowly than divisions, so the count of segments always runs ahead. Divisions settle first because the habit is to prepare the container before the contents. What preparing the container first brings with it is treated in marketing design as a whole.
There is a problem in the axis of division too. The axes presented are almost all attributes. Region, age band, purchase history, occupation. Dividing by attribute contains an inference: applying what was observed about a group to one person belonging to it. A named error in statistics corresponds to this inference. A relation holding at the level of groups sometimes fails at the level of individuals, and can even run the other way.
What follows is treated, and what is not, set side by side. Treated: why the attribute axis gets selected; what dividing by state means; and how to read whether the axis you currently use is a state. All three get answers in the text.
Two things are not treated. How many segments is the right number does not come out here. That value is settled by the kinds of content you hold, and the stock exists only on your side. Nor is how to implement any of this in a delivery tool. Implementation procedures differ by tool, and once the axis is settled, implementation follows.
What the damage looks like can be set out too. Take an inference as a conclusion and the way of correcting narrows to one when it starts to miss. Not knowing what sat between the attribute and the response, the only prescription that emerges is to cut the attribute more finely.
And that prescription can always be carried out. Attributes were captured at registration, so re-cutting requires no new material. A prescription that concludes with material already on hand removes causes outside that material from consideration.
📖 Contents
- The Conditions Offered for Segmenting Are Almost All Attributes
- The Inference Actually Performed When Dividing by Attribute
- What Settles the Axis Is Not Importance but Ease of Capture
- Purchase History Is Near a State, but What It Shows Is a Past State
- The Axis You Divide By Fixes the Recognition of the Divider
- What It Means to Divide by State
- A Fast Way to Learn a State Is to Have the Other Person Declare It
- Increase the Segments and the Operation Breaks Before Anything Else
- Read the Names of the Segments, Not Their Number
The Conditions Offered for Segmenting Are Almost All Attributes
Every piece of the advice on offer holds up as practice. The selection of axis, though, leans towards a single kind.
Segmented delivery means filtering a list by conditions and varying what is sent by target. It also goes by targeted delivery, set against the form where the same thing goes to everybody at once.
The axes offered are attributes and history. Region, age band, sex, occupation, industry, company size. Whether they have purchased, what they bought, when, how much they spent. Combining these into conditions is the form presented.
Four effects are described. The proportion opened rises. The proportion clicked rises. The proportion reaching application rises. The number who stop the delivery falls.
As execution advice, not dividing too finely at the outset is recommended. Start with two or three segments. The reason attached is that preparation grows with each division, so the operation stops being sustainable.
Every piece of this advice is accurate as practice. Content nearer to somebody’s interest does raise response, and unrelated content arriving repeatedly does get stopped. Finer division does increase labour.
The set carries a shared premise. All of it presumes that an attribute stands in for an interest.
Content for thirty-something office workers to thirty-something office workers; a notice about the next product to those with purchase history. The operation of estimating interest from attribute has become the substance of the operation called dividing.
This is not to say the estimate always misses. Fields where attribute and interest are strongly bound do exist.
It breaks when the fact that this was an estimate is forgotten at the point of execution. The moment the condition filters, that group is handled as “people who hold that interest.” That it was an estimate does not survive into what gets sent.
No step for confirming whether the estimate landed is prepared anywhere. If the response rises the estimate landed; if it does not, the division was too coarse. Neither reading ever doubts the attribute axis itself.
The four effects also differ in nature. A rise in the proportion opened and a fall in the number who stop the delivery are two faces of the same operation. Stop unrelated content arriving and both move.
A rise in the proportion reaching application, though, requires a further factor. Nearness of content is insufficient; a need has to have arisen at that moment. Line the four up in a single row and this difference disappears.
That the four move together is common enough. Being common, they have been spoken about as a single effect. When they fail to move, which of the four failed cannot be known without looking at them separately.
The Inference Actually Performed When Dividing by Attribute
The operation of dividing splits into three steps.
First, the list is filtered by a condition. What is obtained is the set of people meeting that condition.
Second, something is assumed about that set. “This layer will be price-sensitive.” “This layer will want case studies.”
Third, content matching the assumption is produced and sent.
The second step is the inference. And the ground for that inference is mostly knowledge about groups. Past data, tendencies shared in the industry, your own experience.
What is being done, then, is applying a tendency observed about a group to each individual belonging to it.
Sociology has long known a caution about this operation. In 1950, W. S. Robinson showed that a correlation calculated at the level of groups and one calculated at the level of individuals do not agree (Robinson, 1950, American Sociological Review, 15(3), 351–357).
The example presented was the proportion of immigrants and the proportion of literacy. Aggregated by state, a relation appears whereby states with a higher share of immigrants have higher literacy. Seen at the individual level, the relation reverses. Immigrants had lower literacy.
Why it reverses: immigrants had gathered in states that were already high in literacy. What was visible at the group level was not a relation between immigration and literacy but a different relation — where immigrants had settled.
Applied to a list, it runs like this. The observation “this attribute layer responds well” does not mean the attribute is producing the response. The possibility remains that people with that attribute happened, for another reason, to be in a situation where responding was easy.
In practice, landing is enough; there is no need to pin down the cause. While it lands, that suffices.
When it starts to miss, what to fix stops being knowable. Not knowing what sat between attribute and response, the only prescription that emerges is to cut the attribute more finely.
None of this means attributes are unusable. Where the state is unknown, an attribute is material for a first estimate.
Read the result of dividing by attribute as “that layer is this kind of person,” though, and the inference has already turned into a conclusion.
The error grows harder to detect as scale increases. With more cases, the group-level figures stabilise. Stable figures look reliable.
What has gained in reliability is the precision of the description about the group, not the validity of applying it to an individual. They are different quantities, displayed as one number on the same screen.
There is no need to land on individuals precisely; a rise in overall response is enough. As delivery efficiency, that is so.
A state where the overall figure has risen, though, contains both those it landed on and those it missed. To those it missed, unrelated content has arrived in the form “this is for you.”
Content arriving in that form is remembered more strongly than content arriving indiscriminately. An error addressed by name outlasts one that is not.
What Settles the Axis Is Not Importance but Ease of Capture
Why is attribute selected as the axis? Not because it best explains interest. Because it is easy to capture.
At the point of registration, attributes are available. Add one input field and both age band and occupation are in hand.
What somebody is currently trying to solve — their state — is not available at registration. Even if it were, it has changed by the following month.
Attributes are static and easy to capture; states are dynamic and hard to. That gap settles the selection of axis.
And once ease settles the axis, the results of dividing follow ease as well. What is on hand is only groups divided by attribute, so the content gets built towards attributes too.
From here one thing follows. Content built towards an attribute can only touch what people with that attribute hold in common. What thirty-something office workers hold in common turns out to be less than imagined.
Because the common holdings are few, the content becomes general. General content lands strongly on nobody.
So the combination arises of dividing finely while the content becomes general. This is not because the division is coarse but because the axis sits somewhere other than interest.
And the combination is hard to detect on your own. What remains on hand is the fact of having divided finely. The number of segments is displayed on screen; there is no field displaying how general the content has become.
The structure whereby ease settles the axis is not confined to attributes. What is measurable becomes the target, and what is not becomes background. This shape recurs wherever indicators are handled. What is peculiar to classification is that what was measured gets attached to people and settles how they are treated thereafter.
It is not that there is no method for capturing the hard-to-capture. States can be captured by asking. Asking costs labour, and some do not respond. So the substance of the difficulty is not “cannot be obtained” but “obtaining costs something,” and the cost is payable.
It nonetheless goes unpaid because the cost is up front and the effect is deferred. Attributes are captured at registration at nearly no cost, and segments can be built on the spot. States require inserting a step of asking, waiting for replies, and reading the distribution before segments can be built. Set the two side by side and the former takes shape faster.
Keep selecting whatever takes shape faster and the slower side is never tried. What is never tried has not been shown to fail; it has simply never had a chance to be compared. That gap does not appear in the list of segments either. What appears in the list are only the segments that were executed.
Purchase History Is Near a State, but What It Shows Is a Past State
Purchase history is a record of behaviour, so it differs from an attribute. This reading is only half right.
Purchase history sits nearer to state. In that sense it is more effective than an attribute.
What history shows, though, is a past state. What somebody was trying to solve at the moment of buying is knowable; what they are trying to solve now is not. And the period right after buying is when the state is likely to have changed.
This lands in the same place as the operation treated in customer lifetime value — estimating the future from a past average. What is on hand is only past records, and what is being handled is a present person.
The period right after purchase is not thereby unusable. Right after buying is also when interest towards you is high.
What should be sent in that period is not necessarily a notice about the next product. The state right after buying is not “finished choosing” but “started using,” and what is needed is not the next set of options but the material for the using stage.
Dividing by history has another gap. Those who have not purchased get pushed into a single group. The “not-yet-purchased” segment. Inside it sit those who have only just heard of you, those who considered and declined, and those with no need. The content requirements of the three differ entirely.
And that segment can occupy a large part of the list. A segment with many people is the one not broken down. Not broken down, the content sent there reverts to “for everybody.” A large undivided segment sits next to the finely divided ones.
History carries a problem of time too. The same purchase count means different things depending on when. Somebody who bought once three years ago and somebody who bought once last month are both “one purchase” in the history. For one the relationship has already ended; for the other it has just begun.
Dividing by count does not hold that difference. Put the date into the condition and it can, and at that moment the segment count multiplies. Where distinct content cannot be produced for the added segments, what increased was the number of conditions, not the classification.
Settings where history does work remain. Where the kind of thing bought directly indicates the problem being solved. There, history reads as a record of state. In a business selling only one kind of thing, this reading is unavailable. With one kind, only the presence or absence of a purchase is recorded.
The Axis You Divide By Fixes the Recognition of the Divider
Dividing has an effect on your side too. The axis you divided by becomes the unit of your subsequent thinking.
Divide a list by age band and the conversation afterwards becomes a conversation about age bands. “The thirties respond well.” “The forties need a different angle.”
Once the unit of conversation is the age band, other ways of cutting stop occurring to anybody. The data on hand is arranged by age band.
This fixing occurs regardless of whether the classification is correct. Use a wrong axis long enough and it becomes the unit of thinking.
Classifications also acquire names. A named group looks as though it has substance. “The beginners.” “The considerers.” “The dormant.”
These are sets of conditions you defined. Once they get called by name, though, people corresponding to those names get treated as real.
Treated as real, accounts of the group’s properties accumulate. “Beginners want concrete procedures.” “Considerers want comparison material.”
The more accounts accumulate, the less reason there is to look at an individual. If the classification is known, the properties follow.
Handling large numbers without classification is impossible; classification is necessary. It is.
The difficulty is not classifying but the classification’s name beginning to function as an explanation.
“They do not respond because they are dormant” is a restatement of the classification wearing the shape of an explanation. Non-responders are what you named dormant, so the sentence explains nothing.
Nor does noticing the restatement settle it. Even knowing it is a restatement, conversation moves faster with a name.
Rename the classification from a state seen from here to a state on the other person’s side, and restatements dressed as explanations decrease. Call it not “the dormant” but “those with no transaction record for six months” and its being a description of an observation survives.
The renaming costs almost nothing. It is a change of wording. And the effect appears directly in the substance of conversation.
“How do we wake the dormant?” and “what is happening to those with no record for six months?” yield different answers. The first searches for methods of acting on people; the second heads towards investigating a state.
Descriptive names, though, run long. Long names get abbreviated in daily conversation and eventually revert to the short original.
Preventing the reversion requires rewriting the segment names in the list itself. The name displayed on screen settles the vocabulary of the conversation.
What It Means to Divide by State
A state is what that person is currently trying to solve.
There are three levels.
The first level is where the problem has not been put into words. Something is not working, but what problem it is has not been pinned down.
The second level is where the problem is in words and methods are being sought. What to solve is known; how to solve it is being compared.
The third level is where a method has been selected and execution is under way. It is progressing, and something is stuck along the way.
The three require entirely different content. Send a comparison of methods to the first and there is no telling what is being compared. Send the articulation of a problem to the third and it is already done.
And the three do not correspond to attributes. People of the same age band, occupation and purchase history may sit at any of the three.
Further, the same person moves among the three over time. Finish solving one problem and the next returns them to the first level.
Dividing by state means dividing by position at a moment, and that position is not fixed.
Here sits the practical difficulty. What is not fixed cannot be stored as a list attribute. It goes stale the moment it is stored.
The method of dividing by state is therefore not storage but declaration each time.
Asking for a declaration costs the other person labour. For that labour, fewer people respond.
And the states of non-responders stay unknown. They get handled as one group, and there, in the end, attribute or history gets used.
The three levels are just your own frame with people fitted into it; that is the same as dividing by attribute. They are similar.
There are two differences. One is that the fitting is done by the other person. With attributes, you filter by condition. With levels, they select for themselves. Where the selector differs, so does whether a miss becomes visible.
The other is that the frame moves. Attributes are fixed at registration; levels can be re-selected. A frame that can be re-selected gets corrected when it is wrong.
There is no ground for the number three. Depending on the field it may be two or four. What is needed is not a number but that the content required at each level genuinely differs. Where it does not, dividing has no meaning. Only the names of the levels increase.
A Fast Way to Learn a State Is to Have the Other Person Declare It
There are two ways to estimate a state. Estimating from behaviour, and having it declared.
Estimating from behaviour needs material. Which pages were viewed, what was clicked, how long was spent. A mechanism for estimating state from these can be built, and being an estimate, it misses.
And the miss is not knowable. When content sent on the basis of an estimate goes unread, whether the estimate missed or the content was poor cannot be separated.
Having it declared is simple. Ask “which applies to you now?” and have them select.
This method has three advantages.
No error of estimation enters. The person themselves is selecting.
Second, the act of selecting is itself an ordering on their side. Selecting among three levels requires thinking once about where they are. The result of that thinking remains on their side as well as yours.
Third, what you handle gets communicated. The way the options are arranged displays the frame you hold.
Declarations are not always accurate either. Some people do not grasp their own state correctly, and somebody genuinely at the first level lacks the vocabulary for selecting.
So the wording of the options matters. Not the names of classifications but the words used by somebody in that state. Write not “the problem-recognition phase” but “it is still not clear what the problem is.”
Asking has a low response rate; most people end up undivided. They do.
Those who did not answer can be handled as one state: “did not answer.” That is an observation, not an estimate.
And the content of those who answered is usable for those who did not. Knowing the distribution of answers to the same question, content can be built for the whole towards the most populated level. It serves not only for dividing but for deciding what to produce.
There is a bias here. Those who answered are a layer with higher interest than those who did not. The distribution is therefore not the distribution of the whole.
The bias cannot be removed. Not removable, it has to be read with the bias accounted for. The shape of the distribution is usable as reference; the absolute proportions are not.
What asking captures is the state at the moment of asking. Three months later it has changed. In a form where you ask once and store the value in the list, the stored value keeps going stale.
If it is stored, when it was asked has to be stored with it. With a date, an old value can be handled as old. Without one, a value from an unknown moment gets used as the present state.
Increase the Segments and the Operation Breaks Before Anything Else
Finer division runs into a ceiling. And the ceiling sits on the operational side rather than the effect side.
Hold two axes and the combinations multiply. Three levels by two purchase states by three interest areas gives eighteen segments.
Producing content for each of eighteen is not actually possible. Not being possible, many of them receive the same content. Receiving the same content, dividing has no meaning.
So where the number divided and the number of kinds actually produced do not match, the classification is a formality.
Further, as segments increase, the record of what was sent to which grows complex. A complex record makes it impossible to judge what worked.
The advice too recommends starting from two or three. That advice is correct, and its stated reason is operational labour.
There is a second reason. The number of divisions cannot exceed the kinds of content you hold.
With only two kinds of content, dividing has meaning only up to two. Divide into three and all you can do is distribute two kinds.
So the order runs backwards. Rather than settling the division and then producing content, the division grows by however much the content grew.
Proceed in that order and the classification starts out very coarse, and coarse classification looks small in effect.
And where the effect is small, the prescription adopted is finer classification. Cut the classification finer without increasing the content and a formal classification is what results.
Increase the segments and the total volume of delivery increases too; prepare a distinct delivery for each and your work multiplies by the number of segments.
More work means less time per message, and the loss appears as density of content.
So the effect of dividing and the quality of a single message compete for the same time. Both can be improved, and doing so requires reducing the total.
There is also the form of sending the same content with only a part varied by segment. In that form, the work does not increase.
What is being varied there, though, is only the address. Where the substance is the same, what is divided is form, not content. Any rise in response is a rise from being named.
The effect of being named is a rise; if it rises, use it. It is usable.
The effect of being named wears out, though, where the substance does not follow. Shown that something is for you, opening it, and finding it was not — repeat that and the showing stops working.
Read the Names of the Segments, Not Their Number
The nature of the axis cannot be read from response proportions.
The first is whether distinct content is actually produced for each segment. Counting settles it immediately. Where the kinds of content are fewer than the segments, the classification is a formality. Where they match, the classification is functioning as a unit of production.
The second: does the same person move between segments? Is the same person in a different segment now than six months ago? Where no movement occurs, the axis being divided by is an attribute. States change, so dividing by state produces movement.
The third: is the segment’s name in a form you could show the other person? “The dormant” and “medium likelihood” cannot be shown. A name that cannot be shown was attached from your own convenience. “Those still sorting out what the problem is” can be shown. A name that can be shown describes a state on the other person’s side.
None of the three looks at response proportions. Since dividing exists to match the other person’s state, what has to be confirmed is the axis.
All three carry conditions of use: several segments, and some elapsed time. Right after starting to divide, there is neither movement nor accumulation.
Right after starting, what can be done is to count the kinds of content you currently hold. That count is the ceiling on how many divisions are possible.
If dividing by state is correct, does attribute data become unnecessary? It does not.
Attributes work as a first estimate towards somebody whose state is unknown. Make the estimate, then ask for a declaration, and the arrangement of the options can be varied.
What can be said is the difference between using an attribute as a starting point and using it as a conclusion. Used as a starting point, it gets corrected later. Used as a conclusion, there is no occasion for correction.
The number of segments sits on the output side of this article’s conclusion. Place it on the input side and the order reverses. Settle the number first and the content thins to match it. Which axis to cut by, and when to ask, can only be settled once the content is settled.
With two kinds of content, only two divisions are possible. That ceiling does not move with the size of the list or the capability of the tool. Of the two figures — the number of segments and the number of kinds — only the second can be moved.
In the end, what your classification is dividing can be read from a single name. Could you show that segment name, as it is, to the people inside it? If you could, what is written there is the other person’s state. If you could not, what is written is a distinction in handling as seen from here, and that is something other than their interest.






