For those who blame the subject line when the open rate falls — What this figure points at is whether the sender sits in the order of priority

What the Email Open Rate Actually Is — It Shifts with the Denominator and Mirrors the Order of Senders

In September 2021, a great many senders saw their open rate rise. It rose without anybody doing anything.

Under the mechanism Apple introduced in iOS 15, mail is pre-loaded by way of Apple’s own servers. Opening is measured by the load of an invisible image embedded in the body, so anything that passed through that route is recorded as opened without a person having opened it. Users who have turned it on are not few.

In other words, what changed the figure was not an operation on your side. The recipient’s environment changed, and the meaning of the measurement changed.

What makes this heavy as an indicator is that the composition of the rise cannot be separated from your side. How far the inflation goes depends on how many users of that environment are on your list, and that proportion shifts month by month. Figures spanning the change are not measuring the same thing. And after the change, the size of the inflation keeps moving as the list turns over.

And the figure has a further property. Among the operations that raise it are ones that merely shrink the list. Shrink the denominator and next month’s figure rises without a word of the content being touched.

What this figure genuinely points at is clear enough. Whether the sender’s name sits inside the recipient’s order of importance. Not the subject line and not the time of delivery. Whether the sender’s name sits inside the recipient’s order of importance is the quantity of involvement they are paying you. Reading that quantity as a stage is handled in what is arranged in the stages.

What gets handled here is the reading of the indicator only. The side of writing is not handled. What gets settled: what the open rate counts; what changes with how the denominator is taken; and how to tell when a fall is good news. All three are answered below.

Two things do not get settled. What percentage counts as good does not come out here. With nobody to compare against, that figure cannot be produced. Nor is when to clean the list. That is settled by the technical conditions of deliverability and by who the business chooses to address, and the second sits outside the indicator.

Where the time goes can be set out too. Chase the figure without knowing it has broken and the time spent on improvement turns out to have been spent chasing a change in measurement. Both the rises and the falls become explicable by factors unrelated to anything done here.

What to write — the side of the text — is treated in writing an email newsletter. Here the indicator is what gets looked at. Until the reading of the indicator settles, the effect of changing the text cannot be judged either.

The Five Methods Offered for Raising It All Genuinely Work

None of the five methods below is wrong. All five, though, stand on the same premise.

The open rate is the proportion of mail delivered that was opened. It is calculated with deliveries as the denominator and openings as the numerator.

The methods offered for raising it run roughly as follows.

First, work on the subject line. Insert appealing words, convey scarcity or a benefit, fit the displayed character count, include the recipient’s name.

Second, the sender display. Set it so that who it arrived from is clear.

Third, the time and day of delivery. Deliver in a band when the other person can read.

Fourth, use of the short preview text. Place a line in the section displayed after the subject line to prompt continuing.

Fifth, narrowing the target. Send only to those with nearer interests and the proportion opened rises.

Alongside these, sector averages are presented, with the form of judging by whether your figure sits above or below.

These are accurate. Subject lines do sway opening, and timing does change visibility at the moment of arrival. Narrow the target and the proportion rises.

Line the five up and the premise sits in a single place. All of them presume that the open rate correctly measures “the proportion of times it was opened.”

That premise did hold, broadly, until recently. The measurement was simple and recipients’ environments were uniform.

What broke is the premise. The breakage is not uniform; its degree varies with the composition of the list. So the size of the breakage cannot be estimated from somebody else’s figure either.

The five show no order. In practice, whatever costs least tends to get touched first. Subject lines are written every time and so touched every time, while narrowing the target requires design. Least cost does not coincide with greatest effect.

Narrowing the target raises the proportion, so surely that is the same kind of operation. It has the same shape.

The difference is where the person who stopped receiving goes. Where the form is to vary what is sent rather than remove them from the list, that person remains. In the removing form, they do not. On the figures, both appear as the same rise.

And the removing form takes fewer steps. Varying the content requires producing as many kinds as there are divisions. Keep selecting the side with fewer steps and the list keeps shrinking. That it shrank does not appear in the proportion, so only the record of having kept selecting remains. On the record, every round went well.

What the Open Rate Actually Counts

What is counted is “the number of times a load occurred,” not “the number of times a person read.”

An invisible small image is embedded in the body of the mail. At the moment that image loads, an opening is recorded.

The gap was small for a long time. The image was loaded only when a person opened the mail.

That premise changed in September 2021. Whatever passed through a pre-load is added as an opening without anybody opening it.

The size of the effect varies with the composition of the list. It is large where users of that environment are numerous and small where they are few. How much inflation actually occurs varies across reports.

Three consequences follow.

First, the absolute value of the open rate cannot be compared with others. Where the composition of the list differs, the same reality produces a different figure.

Second, it cannot be simply compared with your own past figures either. What is being measured changed across the point where the mechanism changed.

Third, the inflated portion arises independently of any improvement to the subject line. Improve the subject line or worsen it and the same portion rides on top.

So is the figure unusable? It is usable.

What is usable is only the trajectory of your own figures measured under matched conditions. Compare across months within a range where the composition of the list does not change greatly. Used that way, the inflation rides on both as a constant, so the direction of the difference is readable.

The composition of a list changes over time. The environments of new arrivals are not necessarily those of earlier ones. There is therefore a ceiling on the period across which comparison holds.

The breakage from pre-loading pushes the figure up. Pre-loaded portions are only added as openings; genuine openings are never subtracted. In environments that block image loading, though, there has long been a separate leak in which an opening goes unrecorded.

As far as pre-loading goes, it does not make the figure in hand lower than the reality.

This asymmetry changes the direction of judgement. When the figure is good, some part of the goodness has to be suspected of coming from measurement. When the figure is bad, the pre-loaded portion does not have to be suspected.

That suspicion about pre-loading runs one way only is a workable property in practice. Check the composition in the months that rose, and in the months that fell look for reasons other than pre-loading. That division reduces the labour of checking.

Why Comparison with Sector Averages Carries No Meaning

Sector averages get presented, with the form of judging quality by whether your figure sits above or below.

The comparison requires at least four premises.

First, that the method of measurement is the same. What can be measured itself differs with the composition of the list.

Second, that the denominator is taken the same way. Whether deliveries sent or deliveries arrived is the denominator changes the value. Whether the undelivered portion is included is a setting on the aggregating side.

Third, that the list was built the same way. A list of only those who applied themselves and a list including those auto-registered through a transaction differ widely in the proportion opened.

Fourth, that the sending frequency is the same. Daily and monthly sending do not give the same meaning to a per-message open rate.

Where even one of the four differs, the comparison does not hold. And published averages often do not state these four conditions.

So what the comparison against an average reveals is only whether your figure sits above or below a published figure. What that signifies does not come out of the comparison.

Without a guide, how can quality be judged? The basis is not outside.

What has to be asked about your own list is “of these people, what share genuinely need this,” and that answer is not contained in somebody else’s average.

Comparison against an average also changes where the target gets set. Once exceeding the average becomes the target, methods of exceeding it get searched for. Among those methods are ones that shrink the list.

None of this means the average itself is wrong. Those who aggregated it produced correct figures about the range they aggregated.

The error occurs in assuming that range is the same as your list. The assumption goes unstated, so the person comparing does not notice they are assuming.

So is somebody else’s figure of no use at all? There is one use.

Confirming the order of magnitude. If your figure is two per cent and every published value sits in the twenties, the magnitude differs. A difference of magnitude is hard to explain by conditions alone. In that case, how the list was built is what comes up for examination.

Where the magnitudes match, nothing beyond that is comparable. The gap between fifteen and twenty per cent is one that differences of condition can account for entirely. Confirming your position inside that band settles nothing about what to do next.

How the Denominator Is Taken Turns the Same Delivery into a Different Figure

This looks technical and is a point that bites in practice.

Take the denominator as “sent” and the undelivered portion enters the base. Addresses that no longer exist, portions refused, portions treated as junk.

Take it as “delivered” and those are excluded. For the same delivery, the latter produces a higher figure.

Defaults differ by delivery tool. And which calculation is in use is sometimes not displayed on screen.

One thing follows. In the month a tool is switched, the figure sometimes moves. What moved is not the contents of the delivery but the method of calculation.

Further, the older the list, the wider the gap between the two, because undeliverable addresses accumulate. With deliveries sent as the denominator, the age of the list appears as a fall in the open rate.

So a fall in the open rate folds together three things of different natures: a problem in the content, degradation of the list, and a change in the calculation.

Separating the three requires watching the undelivered count separately. That number is sometimes displayed on the same screen and attracts far less attention than the open rate.

Which denominator is correct is not the question. It varies with the purpose.

Follow a trajectory without knowing which is in use and a mid-course change of calculation goes unnoticed. Unnoticed, the cause of the change gets attributed to the content.

Confirming is simple. For a single delivery, pull both the sent count and the delivered count off the screen and divide the openings by each. Which one the displayed open rate matches reveals the setting.

The gap is slight, so is it worth the attention? On a newly started list, the gap is slight.

It widens with time. Addresses fallen out of use, addresses changed, portions that stopped receiving — these accumulate.

So the size of the gap is itself an indicator of the list’s age. Once the two calculations start diverging, the state of the list comes up for examination before the content does.

And the undelivered portion is a figure to confirm before the open rate. What has not arrived does not get opened however the subject line is crafted. In sequence, this sits upstream. This has the same shape as the diagnostic order treated in where to start when the conversion rate is poor.

The reason upstream figures attract less attention than downstream ones is that they move little. The undelivered proportion does not change much month to month. A figure that does not change offers no monthly reason to go and look at it.

The Easiest Way to Raise the Open Rate Is to Shrink the Denominator

This indicator has the same hole as any other proportion. Shrink the denominator and it rises.

Concretely: remove from the list those who have not opened for a while and next month’s open rate rises. Not a word of the content moved.

The operation is sometimes recommended in delivery practice. Keeping undeliverable or unresponsive addresses degrades the deliverability of the sending itself — there is a technical reason.

The reason is sound. Where the result of the operation gets recorded as “an improvement in the open rate,” though, what happened stops being knowable.

The reason it stops being knowable is that two operations move the same figure. A case where the content improved and more people opened, and a case where non-openers were removed and the denominator shrank, appear as the same rise.

The way to distinguish them is to look at the absolute count. Where the count of openings increased, the former happened. Where the count is flat or falling while the proportion rises, it is the latter.

And the latter cannot be repeated. There is a limit to who can be removed. Do it twice or three times and the list becomes only the most responsive layer, with the figure stalling high. At the moment it stalls, the number of people reached is a fraction of what it was.

There is no point sending to people who do not respond. There are cases where there is not.

Not responding and having no relationship are different, though. People who never open, remember the name, and search their way back when a need arises do exist.

On the open rate that person is worth zero, and to the business they are not. Tidy a list to fit the indicator and that layer disappears first.

None of this says cleaning should never happen. Deliverability problems are real.

The figures in a month where cleaning happened cannot be used to evaluate content. To use them, the record has to be split before and after the cleaning.

Repeat the cleaning and your own perception shifts too. What remains is only the responsive layer, so every send produces a response. With a response every time, it feels as though what is being written reaches widely.

In fact the range reached has narrowed. That it narrowed does not appear in the proportion. It appears in the absolute count, and the absolute count is sometimes not set as a target.

Writing towards the remaining layer is not itself a mistake. The clearer the target, the more concrete the content.

What is troublesome is that the remaining layer is a result of being trimmed by response rather than a result of choosing an audience. Trim by response and those who need the content but do not signal it in the form of a response fall outside the audience.

Put a Single Indicator Up as a Target and Its Own Kind of Breakage Appears

The three of inflated measurement, the denominator and list cleaning all bite in the same place. They become problems only where the open rate is set up as a target on its own.

How indicators break when set as targets was classified early on the management side. In 1956, V. F. Ridgway sorted the dysfunctions that performance measurement brings to an organisation and held that the breakage differs across three cases: a single measure, several measures set side by side, and several composited into one (Ridgway, 1956, Administrative Science Quarterly, 1(2), 240–247).

What occurs with a single measure is the sacrifice of the unmeasured dimensions. Behaviour gets rearranged so that the one measured figure improves, and the unmeasured parts become the cost. Set only the open rate and the number of people reached, and what remains after opening, move to the cost side.

What occurs with several set side by side is collision between indicators. Set the open rate and the absolute count together and an operation raising one sometimes lowers the other. Without a settled priority, judgement wavers on the spot.

What occurs with a composite is concealment of the weighting. Multiply several figures into a single score and how much each element contributes becomes invisible. Being invisible, operations that raise the score gather on whichever element moves most easily.

The classification does not depend on the size of the business. Even sending alone, if what is measured is one thing, the single-measure breakage appears.

What Ridgway observed were control systems in large organisations. In a one-person business, with nobody standing between indicator and behaviour, collisions surface less readily. In exchange, nobody points at the cost side.

So the available forms for handling the open rate narrow to two. Stop setting it alone and read it paired with the absolute count. Or take it off the target list and place it as a background figure. Both are operations for avoiding the single-measure breakage.

Of the two, the second is easier to carry out. Set as a pair, a judgement is required in any month where the two collide. A form requiring judgement reverts, in periods when judgement is unwelcome, to looking at one side only. Taken off the target list, there is nothing to revert to.

Opening and Finishing Are Separate Events

A rise in the open rate does not mean more got read.

Opening is a record of having opened. Closing after three seconds and reading to the end are the same single case.

And raising the intensity of the subject line sometimes raises opening while lowering completion. Drawn in by a strong subject line, opening, and leaving because the contents differed from expectation.

At that moment, the open rate is recorded as an improvement. What was recorded as an improvement becomes an input lowering the open rate of the next delivery.

So chase the open rate alone and a period arises where the reverse direction goes unnoticed. It gets noticed months later, when the open rate begins to fall.

Detecting that divergence requires looking at what comes after opening. How often a link placed in the same position within the body was clicked. Seen as a ratio against opening, it approximates the share of those who opened and reached the end.

Some ways of reading involve no clicking; is that not inaccurate as a completion indicator? It is inaccurate.

Somebody who read without clicking and somebody who closed without reading cannot be separated. Even so, it measures a place nearer to the content than the open rate does.

The value of measuring a nearer place lies not in the accuracy of the absolute value but in the small number of reasons it moves. Opening moves with the subject line, the timing and the measurement mechanism; whether somebody reached the middle of the body moves mainly with the content.

Nor does a high click proportion mean the content is good. Ways of writing that induce clicks exist, and disappointment past the click echoes in the next round.

The ratio between the two figures mirrors the relation between subject line and content. Neither alone mirrors it.

Following both at once complicates judgement. In practice, placing one as principal and the other as a check is workable.

What should be principal is the one nearer the content. Place opening as a background figure for confirming the name’s position. With that order, the time spent improving subject lines is less likely to swell.

Place a content-adjacent indicator as principal and the figures move less. Opening moves daily; how content lands changes only over months.

Watching a figure that does not move meets resistance, because no sense of doing something arrives. Chasing the small ups and downs of the open rate produces a stronger sense of working.

That difference in feel builds priority in practice. The moving figure becomes principal and the still one gets demoted to a month-end check. The demoted figure eventually stops being looked at. There is no field anywhere for recording that a demotion occurred.

When a Falling Open Rate Is Good News

There are three cases.

First, when the subject line was made accurate. A subject line accurately stating the contents makes it possible to skip an unrelated round. The better people become at correctly skipping, the lower the open rate goes.

Here the proportion opening falls and the proportion, among those who opened, who needed the content rises. The indicator deteriorates and the reality improves.

Second, when the list is growing. New arrivals do not yet hold the sender’s name inside their order. The more people arriving, the lower the overall open rate.

Split the older and newer layers of the list here and it becomes apparent that nothing has fallen. Look at the whole undivided and growth appears as deterioration.

Third, when the frequency was raised. The per-message open rate falls while the total volume reaching people rises. Which one is watched reverses the evaluation.

Common to all three: the open rate on its own does not indicate direction. Whether a fall is good or bad cannot be settled without knowing what else was happening at the same time.

This property causes trouble in the setting of reporting. Where figures get lined up, a fall demands explanation. Explanation being troublesome, operations that produce falls get avoided.

Among the avoided operations are making the subject line accurate and growing the list. Both are correct operations for a business.

Reading every fall as good news is the opposite error. Content having thinned is in fact also a cause of a fall.

What judgement requires is a record of what was done in the same period. Whether the form of the subject line changed, whether the list grew, whether the frequency changed. Without that record, the figures cannot be interpreted.

The record fits in one line. Beside the date of the delivery, write what changed in that round. Rounds where nothing changed leave a blank.

A run of blanks is a period usable for comparison. Insert a round where something changed and the two sides cannot be compared directly.

Start keeping this record and months carrying several changes at once may turn up. The form of the subject line changed, the frequency changed, the list was cleaned. Nothing can be read from the figures of a month where three things moved together.

A run of unreadable months is itself a determination. A state of continuous improvement where what worked is unknown shows that the speed of changing exceeds the speed of measuring. As long as it exceeds, the total volume of improvement grows while what is known does not.

Split the Proportion Three Ways and See Which Part Moved

The proportion itself does not indicate direction. Whether the numerator or the denominator moved is written only outside the proportion.

The first is the trajectory of the absolute count of openings. Line up counts rather than proportions. Where the proportion rises while the count stays flat, what moved is the denominator. Where the count grows, more people are being reached. Watching proportions alone leaves these two mixed.

The second is the open rate split by the period in which people joined the list. The layer that joined three months ago, a year ago, before that. Where older layers are lower, the name’s position is slipping. Where every layer is the same, the overall variation is being produced by the composition of the list.

The third is how often a link in the same position in the middle of the body is clicked relative to openings. Where openings rise and this ratio falls, the subject line has become stronger than the content. Where the ratio holds while openings rise, the name’s position is rising.

None of the three looks at sector averages. Since opening is a manifestation of the name’s position, what has to be confirmed is the inside of your own list.

The three carry limits. All of them require several deliveries and some number of cases. While the count is small, the proportion swings widely. Send to ten and three open and it is thirty per cent; four and it is forty. One person moved.

What to look at at that stage is not the proportion but who opened. At a scale where individuals are legible, individual records carry more information than aggregates.

The second indicator cannot be used without recording the period of joining. That worry can prove groundless.

The registration date is often retained automatically, because many delivery tools record the timestamp of the application. What is missing is the operation of splitting and aggregating by that date.

The operation goes undone because it is tedious, and the material is already in hand. Split it once and what the overall figure was an average of becomes apparent.

This article holds no figure for a standard. The reason it cannot is straightforward: where the way the list was built, the material being handled, and the time since delivery began all differ, the same figure points at a different state. The timing of cleaning and the frequency go unsettled for the same reason. And hand over a figure that does not apply, and the person receiving it tidies their list towards it. The result of that tidying is one of the three breakages this article has been treating.

What changed the figure in hand was not something done here. Something that happened outside changed the meaning of the measurement. The same thing happened in the month a tool was switched and in the month a list was cleaned. Look at a figure that moved and search for the operation that moved it, and something is found. What is found is not necessarily the cause.

If so, what needs adding is not an indicator but a record. Beside the date, write one line about what changed in that round. The form of the subject line, the frequency, the growth or shrinkage of the list, the settings of the tool. Rounds where nothing changed may be left blank. Where two blanks run together is the first interval usable for comparison. Until then, no figure is a candidate for interpretation.

上部へスクロール