Whether trying to increase user goal completion or long-term user trust, we've had to get comfortable with the fact that being data-driven isn't enough, and that doing the best thing for users isn't free.
Science and Sensibility: Thoughts on Experimentation and Growth
























Auto-generated transcript - may contain errors. Tap a timestamp to jump the video.
So good afternoon. Thanks for sticking with us. It's a big day today, loads of talks. Skyscanner started about thirteen years ago as three guys in an attic in Leith with a spreadsheet that they use to find the cheapest flights to take their ski holidays every year.
And thirteen years later, we do effectively the same thing. We just do it at scale. So we aggregate millions of prices and content from thousands of travel providers. And when you come to our site, we help you compare all those options and find the best flight for your trip.
Now I joined Skyscanner just over three years ago, and so I've had a chance to see some of that scale in action and some of the differences that it's made in how we approach product development. So today, I just wanna share a few thoughts with you about the role of experimentation in particular and go through three case studies, which show some of the things we're learning from experiments.
So first and I was worried about this clicker, so we're just gonna we're gonna do a test here. K. So first things, in two thousand eleven, Eric Reeves published his book, The Lean Startup, and since then, it's become best practice for how we do product development in startups and and many other Internet companies as well.
And the basic premise of the Lean Startup is that we need to identify our hypotheses about who our customers are, what they need, and we need to validate each of those hypotheses in turn using an experiment. And in his book, Eric Rees outlines this in three stages, the build, measure, learn loop, and we typically tackle it in reverse order.
So we think about which assumption or hypothesis we want to learn about, and then we figure out how we can measure whether or not that hypothesis is true, and then we build an experiment that helps us take that measurement. And finally, we take the learning from that measurement back into our next iteration of the loop.
And two more critical things about this methodology. So one, the whole point of it is to reduce waste in our product development cycle and make sure we build something that somebody genuinely needs. And the second thing is, we wanna go through that loop as quickly as possible because that's our engine for growth.
If we can quickly learn about what our customers need and equally what they don't, then we can drive accelerated growth in our company. Now if you think about those things, it's called the lean startup methodology, but big companies want that too. Right? We want to eliminate waste in our product development cycle, and we want to have the growth profile of a startup.
And the the thing is that some of these dynamics change as you scale and lean start ups doesn't apply in the same way. So just to give you some figures on how Skyscan has grown, we now have fifty million users a month. We have eight hundred people in our company, and we have forty product managers in the team.
And some of the things that are different when we're trying to apply the lean startup methodology at this scale is that for one thing, when you have that many people, you have to divide them into teams. Right? And they can't all focus on the same thing.
They can't all focus on the major metric for your company. They have to start to optimize in different areas of the product. So that's one big difference from a start up to a larger organization. The second thing is you're just focusing on different problems.
Right? When you're a start up, you're trying to figure out who's your first customer, what value, can you bring to them? When you start to scale, you've largely answered those questions already. So the kinds of problems you can use to continue to drive growth are very different.
And the third thing that's really different is just quite practical. As a start up, you have to be laser focused because if you introduce waste into your development cycle or you build things that people don't need, then you go out of business. Right?
As you scale, start to have a little bit more of a cushion, and so you can relax a little bit and not go out of business tomorrow. There's always a risk that a competitor will eclipse you a year down the line, or some other startup will come up and make you redundant, but that risk is way more remote for your workforce as you start to scale.
So these are some practical differences when you're trying to apply the lean startup as your company grows. And the risk then is that as you go from start up to scale, your lean your build, measure, learn loop starts to look like like this.
Instead of being a really quick, tight iteration, it starts to kind of slow down and be a bigger iteration like this. And instead, what we really want is to have lots of tight iterations like this, and we want to take advantage of the fact that we have lots of people, lots of teams, we want to be running build, measure, learn loops in parallel.
We want them to have the intensity of focus that they have in a start up so that we can have the same growth profile. So I want to talk to you now about something we've been doing at Skyscanner. We've had a lot of practice recently thinking about these loops and how to eliminate waste in our development cycle because we've run hundreds of them.
And one thing, one advantage of having scale in lots of users is that you can run-in particular a b test, and that's just one form of experiment you could use in a build, measure, learn loop, but it's one that's particularly effective at scale.
And what an AB test is is you you give two variants of your product simultaneously to two different cohorts of your user base. And by doing that simultaneously, the main thing you get from an AB test that you can't get anywhere else is you can identify the causal relationship between the change you've made to the product and the impact it's having on your users.
So we decided that we really wanted to start doing AB testing at scale, and about eighteen months ago, we'd never run a single AB test, and this chart shows our acceleration. And last month, we ran about two fifty tests. Each of these was an opportunity for us to practice, build, measure, learn at scale.
Right? So I'm going show you one other data point which is pretty interesting, which is the count of individual hypotheses that we were testing with these experiments. And this isn't a perfect metric, we're kind of collecting it manually. But even if it's not perfect, I think it's pretty obvious there's a big gulf between these two lines.
Right? So what this shows is that for every hypothesis we want to test, we're running multiple experiments, and that looks like a form of waste in our product development cycle. Okay. So what's going on there? So let's look at a hypothesis first in more detail.
This is a standard format we use for all hypotheses at Skyscanner, and I know we saw another example from a talk earlier today. This is pretty similar. So based on a particular insight, we believe or predict that a specific product change can cause a very specific impact.
An example might be, we observed that travelers disproportionately purchase flights to very popular destinations for their holidays, and so we believe that by changing the sorting of our results to popularity instead of price, we can increase conversion rates. That would be an example of a hypothesis.
And with a hypothesis like that, there's two answers, right? Either you run the test and it's successful, so we change the sort order and indeed we do see an increase in conversion rate, or we run that test and we decrease conversion rate, and from either of those, we learn something we can take into the next iteration of our build, measure, learn loop.
Now if these are the only outcomes that we had at Skyscanner, we should see a one to one relationship between the two charts I showed you earlier. Right? For every hypothesis, we run run run one experiment and we get an answer. So that's obviously not happening.
So we looked at some of the tests we were running, and we found that there were two other outcomes. And as of four, I'm glad the slide came with me. Okay. So first, successful and failed, and then two other outcomes, invalid and flailed.
So an invalid experiment is one that looks a lot like like a successful experiment, except that it's successful because it could not have failed. And we'll talk a little bit more in one of the examples about how that happens. The second example is flailed, and a flailed experiment is one that failed because it never could have succeeded.
And I have a yeah, this guy. This is how you feel when you run one of these experiments, and that's why it's called a flailed experiment. So I'll give you an example. Okay. So the first case study, we ran an AB test where we have this insight that we have a flights product, a lot of people like us for flights comparison.
We also have a hotels product, and we do exactly the same thing for hotels that we do for flights. We aggregate hundreds of results, millions of results, and we allow you to do comparison on our site to find the best deal for your trip.
And so our insight was, if people like us for flights, they probably like us for hotels, if they like that value proposition. And also, it's pretty obvious that lots of people who buy flights also buy hotels for those same trips. So we ran this AB test, where instead of redirecting people or landing people on our flights homepage, which is on the left, we landed them on our hotels home page.
And the idea was, if we introduce our users to hotels, what we should see, if we're successful, is more searches on our hotels product and more conversions on our hotels product. Okay. So what happened? Well, we ran this test and what we saw was indeed many more searches in hotels and many more conversions.
And we also didn't notice any degradation in the flights metrics, so we saw that even though not everybody obviously wanted a hotel when they landed on our website, they found it pretty easy to navigate back to flights and continue with their journey. So that looked really successful.
Well, it looked like that until we started to get some really interesting user comments. And I'll show you this one. These are amazing. Right? This guy is so angry that we've run this test. And he writes, Skyscanner, you are really annoying me with your hotels and car hire options.
You are not called a hotel scanner or car scanner, and I can't I can't argue with that. Right? And if you want to diversify into other services, you should have thought of that before you named your company. And then he gives this amazing example of compare the market, and if they'd chosen to be called compare the car insurance market.
And then finally, he ends by saying, please stop defaulting me to hotels. I'm here for flights. And this is this example is so interesting because you can see he's he's clawing at all the logic he can come to to try and compel us not to run this test anymore and to change the behavior.
And so we call this test, in the end, invalid. And the reason is it was really, it was just not capable of failing during the a b test. We were able to measure the upside, which was the increased conversion rate in searches in hotels, but we weren't in the a b test able to measure the agony that it was causing some of our users and the friction that inevitably would have left them, led them to leave our product.
So that's an example of an invalid test. And the hallmark of these tests, the thing to watch out for in your own organizations, is that they tend to happen when you're trying to release a new product or a new feature, and you're so concerned about, seeing whether people engage with that feature that you don't develop an experiment that, actually tests the possible downside.
That's how they tend to occur. Okay. So next case study. We wanted to increase the propensity of travelers to enter our search funnel from our news pages. Now we write lots and lots of travel content, and that content gets picked up by news organizations, and it gets pushed around social media, and millions of users arrive on our website through this kind of content.
And the thing that's interesting about that for us is that the moment for our customers is when they do their first search. That's when they understand the value that Skyscanner provides above and beyond going to an airline site or a travel agent, it's when they do a search.
And so here we have millions of users landing on our site. They tend to read these articles and then leave again. And so we wanted to run a test to see if we could bring more of them into the search funnel. So let's see. Yeah. Okay. So now here's the test.
We put our search controls, which is that big gray box on the right hand side. We we pop them onto the news pages, and the hypothesis here was these controls work very well. They're highly optimized. They're on our homepage, and they're a trivial way for us to introduce travelers to search or sorry, readers to search.
And also, we know that these people are interested in travel because they're reading our travel content. Okay. So what happened? Well, we ran this test, and nothing happened. So we put these search controls on on these pages, and you cannot miss them. Right?
But no one was entering the search funnel or it was just a negligible amount. And so our verdict here was that this test was flailed. It could not possibly have succeeded and it's kind of easy to see that looking back. The people reading these pages, they don't necessarily know what Skyscanner is yet.
Right? They've just found a link on a on a social site. They found it on a news site. They wanted to read this content, but they're not necessarily looking to book a flight today, and they don't know that that's what we offer. So this is a really clunky way of trying to introduce them to the value proposition that we have.
And this test, in the end, probably wasn't worth running. We should have just thought this through in advance and then thought of a way to do a more targeted introduction to our product for this type of user. So again, the hallmark of a test like this is where you, where you design something, sorry, where you design something which is probably a little bit too minimum to accomplish your goals.
So you, the conversations that will start leading you to a flail test are things like this. Why don't we just test this idea? Or I don't know. Why don't we just throw it out there and see what happens? That should be like a big red flag that you're about to do something that is totally worthless.
Okay? So final case study, the right metric for travelers. Earlier this year, I was writing one of the most difficult emails of my life to our senior management team, and the email was basically this. We'd found two major problems for travelers on our site, and the good news was we found a solution, but it required voluntarily taking the single largest degradation to our conversion metric that we had ever seen.
So it was not a pleasant email to write. And I'll just show you what the problems were. So here's a traveler on our site, and we we started to see these usability tests. They would find a flight, and then they would have this thought, okay, I like this flight. Now I wanna see more detailed information.
And so we have a very big call to action button, it's bright green like it shows in this cartoon, and they would tap that button. And immediately, they would be redirected off our site to book with one of our partners. And the reason is we think if you click that button, you're ready to go book.
And so that's what we do. We send you off to the partner site to make a booking, and we count that as success because that's our number one conversion point in our company. Right? So here we identified a case where we were hurling people off our site who were not ready to go, we hadn't solved a problem for them, and we were counting it as success.
That was problem number one. Problem number two, we identified in our analytics, which was that travelers appeared to be only really choosing the cheapest provider that we have on the site. And this was important to us because our whole job is comparison, and if everybody just chooses the cheapest provider, it's not obvious that travelers are really engaging in comparison on our site, and so maybe they're missing part of the value we could be bringing to them.
So we'd identified those two problems, and then we ran a test to try and fix it. So here we are again, and I'll show you the new layout that we tried. So same user finds a flight that they like, and this time when they hit the select button, we changed the layout.
And when they hit the button, instead of immediately ejecting them from the site, we dropped down a list of all the providers who they could choose to book that flight with and the prices. Okay. Now this doesn't answer the traveler's problem immediately, but suddenly she's still on our site and she realizes that was not the button I was looking for.
Right? And then she notices we have a details button. And when she clicks that button, it takes her to, the content that she's looking for. In this case, it was, more detailed information about the stopover, how long it is, etcetera. Yeah, I think someone at the back is helping me because I'm just really struggling with the clicker.
So perfect, we solved her problem. What about the other problem, that people aren't doing comparison on our site? Well, we also found that by running this test, people started to look at providers lower down the list, and in fact, just by changing this layout in this way, we increased comparison behavior by four hundred percent.
I'll show you that now. Here's a guy saying, Hey, this flight looks good to me. He clicks the select button and we drop down that list of providers, and now he realizes there's somebody else he could book with like British Airways and he's got a loyalty scheme with them, so he chooses them instead of the cheapest provider.
So that's success on both counts. Two major problems solved for travelers. We're increasing engagement, with travelers on our site, and they're getting more value out of our product. Right? But in order to do that, we have to stop ejecting them off our site when they're not ready.
And the trouble is that's how we count conversion at Skyscanner. So in the end, we called this test successful, but it was because we discovered there was an underlying flaw in the way we were counting success. Right? Because normally, if you introduce a major drop in conversion rate, that would be a pretty easy test to call failed.
In this case, the value for travelers was still there, so we called the test successful, we've implemented this change now in our product. Okay. So what have we learned? Using lean startup methodologies, is meant to help us reduce waste in our product management cycle.
But when you start to scale, those dynamics change. And we've walked through two case studies today of, things you can look for in your own experimentation where you might be introducing waste into your product development cycle, and those are invalid and flailed tests.
We also looked at an example where it would have been very easy to make the wrong decision about what was best to do for travelers because the metric itself had underlying biases and assumptions. And so what we've learned is that it's really important when you're building a data driven culture to balance both science and sensibility.
Eric Rees wrote in his book something reasonably prescient, which was that anytime a team attempts to justify its failures by resorting to learning as an excuse, it's engaged in pseudoscience. And we find that as our company scales, we're perhaps more susceptible to engaging in pseudoscience because we don't have that concrete risk of going out of business like a start up does.
It's very easy to engage in lots of experimentation, not have a direct benefit for travelers. So as we learn about these things, we're writing a lot about what we learn and we're also open sourcing some tools on our engineering blog, Code Voyagers, and there's two specific things we've open sourced, which may be of interest to you.
We have a template for writing hypotheses and for experiment plans. We also have an analysis toolkit for AB testing. So I invite you to well, go see what we have and use them if it's useful to you. So my final thought for you today is that as we scale, we have to be on our guard against this kind of pseudoscience, and the only way around it is to apply both science and sensibility.
And as product managers and product people, you have a critical role to play in helping your organizations find that balance. Thank you very much.