The Case of the Royal Wedding

This is a story that gives you a sense of some of the problems that we had in the early days of Twitter Ads. It’s based on a presentation I gave for Twitter’s “War Stories” series in 2015. I included screenshots and charts from that presentation that have more internal details than I would normally share since this happened 15 years ago and Twitter as a company no longer exists.

It was the spring of 2011 and love was in the air. And not just any type of love: royal love. The entire world was preparing for Prince William’s marriage to Miss Kate Middleton. Buckingham Palace had issued a 91-page press release with details about the motorcade, the wardrobe, the ceremony. On page 56 they wrote, “Twitter has said it may need to bring in extra servers to cope with the anticipated demand on that day.” This is just one more exhibit in the long history of the gap between royalty and reality; we said no such thing. We did, however, bring in one extra server just to be safe.

This was a different time at Twitter. We only showed ads in search results. Ads were only on twitter.com, not in our mobile clients. Promoted trends were our biggest source of ad revenue; they brought in a modest $30 million a year. There was only one worldwide promoted trend each day. But not everything was different. In 2011, we also wanted to make money. So we took the royal wedding and sold #RoyalWedding (Promoted) to SlimFast.

Photo of Prince William and Kate Middleton waving with the superimposed text "#royalwedding Promoted"

When you bought a promoted trend you were actually buying two things: placement at the top of the trending topics, and exclusive placement in search results for that trend. Exclusive means that you paid a flat rate and didn’t compete in an auction. Normally, the ad server runs an auction every time a users does a search to determine which ads to show. The winner is determined by advertisers’ bids and the predicted click-through rate. (The winner pays the runner-up’s bid for each engagement; it’s a Vickrey auction. It’s also more complicated than I remember now or understood then). To get your ads to show up, you pay more, have more engaging content, or buy exclusive placement and skip the auction. This way when someone clicked your trending topic, your ads were guaranteed to appear at the top of the results.

On the engineering side of ads, we had just finished this big migration out of our monolithic Rails codebase. Following that we had just solved a doozy of an incident where sometimes an advertiser would sign in and be another advertiser, see that advertiser’s data, and change it. After those sprints we were feeling winded.

About two weeks after the wrong-user incident, on a Wednesday afternoon, I was busy working on the upcoming launch of a product called Promoted Tweets in the Timeline when I got an email from one of my colleagues on the sales team in New York, subject: “VZ NFL Draft Talk at 3pm EST?” This was not uncommon; from time to time on the engineering side, we’d get questions from our coworkers in sales. We went back and forth a little bit about how to configure an upcoming campaign for Verizon, who was going to be advertising on the term #NFLDraft.

A few messages later I thought we were wrapping up when I got this reply:

From: Sales
Subject: Re: VZ NFL Draft Talk at 3pm EST?
Date: Wed, Apr 27, 2011 at 4:56 PM Hey Ryan Quick question, will the spend be taken from the promoted tweet campaign or should we create a separate IO in the system since this is supposed to be a flat rate?

Flat rate? Promoted trends were the only flat rate/exclusive product that Twitter had, and we had already sold #RoyalWedding to SlimFast for tomorrow. What’s going on? Utkarsh, who was the tech lead for ads engineering at the time, jumped on a phone call with the account exec (in corporate America at that time you never made a phone call, you always jumped on one). He reported back with a bit of a quandary:

From: Utkarsh
Subject: Re: VZ NFL Draft Talk at 3pm EST?
Date: Wed, Apr 27, 2011 at 5:23 PM … The big issue is that these campaigns are not guaranteed. That will enable others to win the #NFLDraft auction. Does anyone know how to fix this?

The problem was that Verizon was expecting an exclusive buyout of #NFLDraft, but we already sold that day’s only exclusive product to SlimFast. Letting other advertisers win the auction for this term would be a no-no. Verizon didn’t spend $100,000 for other advertisers to show ads on their keyword. It would be like booking a Super Bowl commercial slot and then giving someone else the air time.

As we chatted about the problem we got the last email of the day from sales:

From: Sales
Subject: Re: VZ NFL Draft Talk at 3pm EST?
Date: Wed, Apr 27, 2011 at 6:31 PM We sold this unit in as a 24-hour roadblock and the client is expecting this to go live at midnight PST.

To recap:

  • We sold something to Verizon for $100,000.
  • It’s supposed to go live at midnight.
  • It’s 6:31 PM.
  • The thing that we sold did not exist.

We needed to create a new exclusive product that was not a promoted trend. We had to change the ads UI so that Verizon could sign in, manage this campaign, and see their analytics. We also had to set up a campaign in the database so that the ad server would load it and serve it exclusively without running a promoted trend.

Andrew, Ben, Colin, Lennon, Tom, and I were still at the office. Andrew and Colin started working on the UI, and then the rest of us dove in to figure out how to set up this campaign. We didn’t have any existing tools since the product didn’t exist. In fact, the business rules for our application explicitly prohibited a campaign with this particular type of setup! So we put on our cowboy boots, saddled up, and went right to the database.

In April 2011, all advertiser data was stored in Twitter’s MySQL users cluster. Tons of services needed to access the database so our DBA team set up a bunch of replicas. In this configuration you have a primary database that handles all of the writes while replicas handle read queries. Analytics had their set of read replicas, anti-spam had theirs, and the ad server had its set. Segregating traffic this way keeps the primary DB responsive while letting services do whatever heavy read queries they want.

This setup worked pretty well, but occasionally there would be a disaster. If somebody somewhere in the organization ran a really expensive query against the primary, or if there was some kind of rogue batch job that was unthrottled and was sending a high volume of writes to the primary, you would get something called replication lag. This is the delay between writing data to the primary DB and those changes being visible in the replica DBs. In a well-running system replication lag is very close to 0. It’s almost always less than a second, and often as low as 10 milliseconds. Unfortunately, on that day the DBA team was having an incident of their own. At the time we learned about the new product we had to build, replication lag in the users cluster was at 5,000 seconds. That meant that after writing to the primary DB you had to wait almost 90 minutes to read it on your replica.

To create a new campaign we had to write to the primary DB, but the ad server read from replicas so after every change we had to wait at least an hour to see if it worked. We opened a production console, turned off business rule validations, and created a test exclusive campaign without a promoted trend to make sure we got the right bits set up. When you are doing risky work like this you want a second set of eyes, a practice commonly called “pairing”. We were quadding on this problem. Every statement was checked, double-checked, then triple-checked by four people sitting around one computer. This seating arrangement generally indicates that something bad is happening.

We created an ad campaign with a test advertiser, @TheSandwichBar, targeting a keyword that nobody would be searching for, #FooFooFoo. We committed the campaign to the database and then waited for 90 minutes.

On a side note, Jamie Oliver, the celebrity chef, was visiting the Twitter office that evening, teaching people how to cook. I had invited my wife to see him, and then almost immediately abandoned her. Our desks were right next to the lunchroom. Jamie was about 50 feet away talking about the evils of junk food and how processed food is killing us. As we waited for the DB write to replicate, we sent Tom to the kitchen to steal a pizza. We scarfed it down while Jamie droned on in the background.

Eventually the ad server picked up the campaign. We searched for the ad on twitter.com, and boom: nothing. What went wrong? At this point, it was 10:00 PM. With 90 minutes of replication lag, we only had 30 minutes to figure out the problem before we actually had to set up Verizon’s real campaign. We reviewed the database queries again. Was the data wrong? We re-read the ad server code. Had we missed something important? Everything looked right, it was just…not working. We racked our brains and then we remembered: we had a product rule that we only showed ads when there was organic content as well. Of course nobody was tweeting about #FooFooFoo. So Colin just sent a tweet. And then boom, this time for real. There it was.

Now that we knew how to do it, we repeated the process with Verizon’s real campaign. On the UI side, Colin and Andrew put up a PR, “Exclusive IO for non-traditional exclusive campaigns”. The change even included tests. At 12:18 AM we sent out a victory email: “The campaign is live. We’re going to deploy the changes to the UI in the morning.” (Deploys on Friday don’t bother me but 12:18 AM does feel like pushing your luck.) We went home, crisis averted.

The next morning, I got into the office early. I had been at my computer for a few minutes when I got a phone call from sales. This was a first for me, but so was the change we pulled off overnight. What I expected to hear was thank you so much; you saved the day; you’re amazing; you’re so technically capable; I can’t believe that you guys were working so late and saved our asses fueled only by ultra-processed foods. The actual conversation, however, was this:

Sales: Hey, what’s going on with the definition box?
Me: The definition box?
Sales: Yeah, the definition box. We sold that as part of the package to Verizon.
Me: We sold what?

On that version of twitter.com there was a cutesy little dictionary definition that told users about Twitter for iPhone, Twitter for iPad, or cool third-party apps. Apparently, this box was part of the mocks that we shared with the client when we pitched them on buying #NFLDraft. Fortunately, there was an internal tool that let people edit the definitions. Unfortunately, I was on the phone because somebody tried the tool and it didn’t work. The reason, of course, was replication lag. At that point it was up to two hours. So it was going to take a little while for the change to show up on our real-time communications platform.

Even after replication delivered this precious database object to all the DBs, we cached the hell out of this data. The definitions basically never changed. As I investigated I realized the data generated by the tool was slightly wrong—it was never going to show up. Back to the production console I went.

For peak performance, we didn’t just cache the DB rows, we also cached the rendered HTML. My plan was to read that value from cache, change it manually, and re-insert it. Step one was to figure out the cache key for an existing dictionary HTML fragment. Then I grabbed the HTML from memcached. I can honestly say this was the only time that I remember pairing on a fragment of HTML. My colleague Avi helped me double-check the changes, then I pasted the HTML back into the production console, hit return, and everywhere, instantly, the definition box was Verizon’s. I like to think of this as the memcache SET heard around the world.

And it kind of was. Later that day the communications team was asked for comment on an article with the headline “Twitter Introduces Text Ads”.

In addition to the non-traditional exclusive campaign, in our last-minute scramble, we’d actually launched a second new ads product that hadn’t existed the day before, Twitter Text Ads.

When things settled down, we did a postmortem and the root problem was just a boring case of miscommunication. The folks in sales believed the product already existed. At the executive level everyone knew about the deal. The key details just didn’t make it down to my team until late in the process. And, in defense of the sales team, the thing that made the product not exist was subtle. Kevin Weil, who was the head of Ads Product at the time, sent out an email saying, “Thanks everybody for scrambling, let’s never do this again.” To my knowledge we never did.

Two weeks later, we finished our database migration that moved advertiser data out of the users cluster (goodbye replication lag). Kevin also sent a note to Tom and to me, “Hey, hate to make you guys look at the definition box again, but let’s kill our new text ad product.”

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论