How I started my testing journey

I had a big problem.

The online fundraising campaign I was running had lost steam. Instead of generating thousands of dollars a day, it was generating just hundreds.

Traffic to the website, driven largely by advertising outside of our control, had fallen off sharply. But the expectations about how much money we’d make from the site hadn’t changed at all.

In short, we had to make more money from fewer website visitors. I knew something needed to change.

My first instinct was to test different elements on the page. With a little skill and a little luck, we would find a technique that would make more money and offset the decline in traffic.

Some of our early tests showed promise. For example, we found that we could boost revenue by changing the gift amount we asked for. We also found that adding an email signup option would (indirectly) get more people to donate.

But none of this made up for the decline in visitors and revenue, which was only accelerating.

The stupidest idea I had ever heard

At this point, my friend and marketing co-conspirator Tim Kachuriak came to me with the stupidest idea I had ever heard. He had just been to a conference, and had really drunk the Kool-Aid. He was extremely excited, and that made me nervous.

His big idea was to test the existing page against a completely new version of the page:

  • Instead of putting the donation form right at the top of the page where potential donors could find it, he’d bury it at the bottom of the page.
  • Instead of including a short-and-sweet paragraph making the case for giving, he’d make page visitors wade through 700 words of copy, some of it repetitive.
  • Instead of using a page layout and color scheme that matched our branding, he’d use a bland, tan-colored page with just our logo on the top.
  • Instead of leading with the website name, to allow visitors to know where they were, he wanted a big long headline.

This “radical redesign” violated just about every fundraising and website best practice on the books.

Not only that, we’d be testing a huge number of changes at once. How would we know what caused the improvement?

I immediately told him he was a fool.

Tim persisted. He explained his rationale for the changes.

The people coming to the site may not yet be convinced that donating is for them, he explained, so we need to sell them before we ask for money. That means we need a lot more copy, to do the explaining. And that means the donation form has to go at the end of all that copy, after someone has bought in.

I wasn’t convinced, but Tim wore me down. I gave in—mostly to get him to shut up about the test already.

Turns out, my gut instinct was totally wrong

It’s a good thing I eventually caved in and allowed Tim to try his harebrained idea.

Tim’s crazy, rule-breaking page performed better than the original. Much better.

It brought in 74 percent more gifts than the existing page. Not only that, the average gift was 179 percent higher.

Combine those together and revenue from the new page was up a whopping 274 percent. That’s more than triple the money.

A new, more robust approach to testing

This result started me thinking about a new, more robust approach to testing.

In our old way of thinking, we tested one page element at a time. We would swap out one image, or one change the text on one button, or modify one headline. This approach had a major upside: because we changed just one thing at a time, we knew that any difference between versions was because of that change. But this approach was also scattershot, with no real rhyme or reason to it; it took a long time to get meaningful results; and we never learned enduring lessons.

In the new way of thinking, we would test not individual elements but rather comprehensive theories of the donor. So instead of testing a red button against a blue button, we would test a page addressed to one sort of donor against a page addressed to a different sort of donor.

This approach has several advantages, as you’ll see as soon as you try it:

  • You learn a lot more. Because you are testing fundamental concepts about our donors, you can extrapolate lessons from one donation page to dozens of others. You understand more about why a page works, not just that it does.
  • You get test results a lot faster. By combining several changes into a single test, you can usually find out awfully quickly whether your concept works or not.
  • You have more fun. In part because you are always challenging the conventional wisdom, running radical tests is awfully invigorating. You aren’t just blindly following someone else’s “best practices” and hoping they’ll work.

So what are you waiting for? Start testing!


Stop screaming

Ryan Levesque:

Marketers still seem to think if they can stand on their tippy toes, and scream loudly enough (metaphorically speaking), even strategically, they can get the attention of customers


So what?

You got a million unique page views last month? You have 5,000 newsletter subsribers? You have 20 percent market share?

Awesome. So what?

If you can’t answer that “so what?” you’re measuring yourself on the wrong thing. You’re using a vanity metric that makes you feel good—it’s a big number—but doesn’t necessarily advance your bottom line.

Worse, since you get what you measure, a bad metric may lead you to invest your time and money where it doesn’t make a difference.

The Periscope team points out that faulty measures lead you to optimize for the wrong thing:

if we were motivated to grow [daily active users], we’d be incentivized to invest in a host of conventional growth hacks, viral mechanics, and marketing to drive up downloads. This direction doesn’t necessarily lead to a better product, or lead to success for Periscopers.

Periscope instead tries to maximize time watched, which they say builds value for both customers and the company.

What would you say if someone asked “so what?” about your metrics?


‘Great website optimizers did not go to school for optimization.’

Brian Massey, quoted by Peep Laja:

Great website optimizers can be found in the most unusual of places, because they are currently among the most unusual of individuals. Great website optimizers did not go to school for optimization. They are grown, not found. They can be found in libraries and in psychology schools. They tend bees, arrange flowers, and any other number of hobbies that involve visual differential analysis.

The whole thing is worth a read if you’re testing online.


Three reasons why you shouldn’t bother with nonprofit benchmark reports

If you wanted to see how your nonprofit’s fundraising stacks up to “industry standards,” you’d likely turn to a benchmark report like the M+R 2015 Benchmark report.

That’s fine if you know what you’re looking at. But nonprofit benchmark reports have several flaws, and if you’re not careful you can draw the wrong conclusions.

1. Benchmark reports are biased

For one thing, the results are biased. I don’t mean that the researchers behind the M+R study and others like it deliberately skew their results. Instead I mean biased in a statistical sense: these reports aren’t based on a random sample of nonprofits. 

Studies like M+R’s draw their data from organizations that volunteered to share their information with the authors. The authors of some studies even examine organizations they have preexisting relationships with, i.e. their clients.

That means most benchmark studies aren’t representative of all nonprofits—which is problematic since the conceit of such studies, their disclaimers aside, is that they are representative.

About all you can say is that the results of the study are true of the organizations studied. That’s interesting but not very helpful to evaluating your own marketing.

2. You can’t draw useful conclusions from benchmark data

Let’s say you work at an animal welfare group. You’re looking to run some Google ads and want to see if there’ll be competition from your peer groups.

Turn to page 38 of the M+R 2015 Benchmark report. You’ll find that 63 percent of animal welfare groups run paid search ads. So you’re behind the curve.

But wait! Is that what the report really says? Let’s take a look at the data.

The M+R report looks at 85 nonprofits, grouped into seven sectors. Of these, only eight are classified as “wildlife and animal welfare.”

What does it mean that there are 85 total nonprofits and just eight animal welfare groups? It means we can’t draw useful conclusions about animal welfare groups (or any other group) from this data. 

Not with any reliability, anyway. For animal welfare groups, the margin of error is ±34 percent at a 95 percent confidence interval.1 For everyone studied, it’s ±10 percent. And this assumes we are dealing with a random sample, which we aren’t. 

So according to M+R’s report, anywhere from 29 to 97 percent of animal groups use paid search ads. Between 48 and 68 percent of all nonprofits do so.

Yeah. That narrows it down.

In fact, you’d need a sample of 370 (randomly sampled) nonprofits just to get to a margin of error of ±5%. That would bring the range down—to between 58 and 68 percent of animal groups.

3. Sector-by-sector breakdowns are meaningless

Nonprofit benchmark studies typically categorize their results into sectors: education nonprofits, internationally-focused nonprofits, environmental nonprofits, and so forth.

This breakdown is useful only insofar as groups that focus on the same set of issues are in any way comparable when it comes to their business and marketing practices. Which is a ridiculous assumption to make.

Can you really compare the ASPCA, which had $166 million in revenue in 2013, to your local animal shelter? By the usual benchmarking taxonomy, you should, because both are animal-welfare groups. 

Or let’s think about the retail world for a minute. If you ran the Apple Store, would you benchmark your sales results against other technology stores, like Best Buy? This is a really odd comparison to make, as this table shows:

Retailer Sales/​square foot
Apple Store $4,551
Tiffany's $3,043
Coach $1,532
Best Buy $808

Sources: Forbes/​eMarketer; Seeking Alpha

The Apple Store isn’t even in the same league as Best Buy, given sales per square foot, a common measure of retail performance. Yet a simplistic definition of sectors, like that used by nonprofit benchmark studies, would lead you believe they’re peers.

So make sure that when you’re comparing yourself to your peers that they really are your peers. Benchmark reports may not give you the data you need to do that.

Why are you measuring yourself against other nonprofits at all?

Charities are notoriously basket cases, with poor management and bad practices. Because they don’t face the same market pressures as for-profit industries, these bad practices can continue for years before anyone fixes them.

So don’t measure your success as a fundraiser against biased data in benchmark studies. Instead, measure your success against your goals: are you meeting revenue targets, retaining enough donors, and acquiring enough donors?

I welcome your feedback. Tell me what you think in the comments.


1. In simple terms, margin of error measures how precise your data is. A bigger sample size will typically reduce your margin of error, meaning your data is more precise.

Let’s say a poll reports that a politician is favored by 50 percent of voters, with a margin of error of ±3 percent. That means his actual popularity (if you were somehow able to talk to every single voter instead of a random sample) is somewhere between 47 and 53 percent. The data captured can’t pinpoint it exactly.

A margin of error of ±34 percent means that your sample data is basically meaningless at predicting the actual data.