# 解剖Twitter 【4】抗洪需要隔离

- 作者：邓侃
- 发布时间：2010-03-09 15:25
- 原文：https://www.ifanr.com/18820

如果说如何巧用Cache是Twitter的一大看点，那么另一大看点是它的消息队列(Message Queue)。为什么要使用消息队列？\[14\]的解释是“隔离用户请求与相关操作，以便烫平流量高峰 (Move operations out of the synchronous request cycle, amortize load over time)”。

为了理解这段话的意思，不妨来看一个实例。2009年1月20日星期二，美国总统Barack Obama就职并发表演说。作为美国历史上第一位黑人总统，Obama的就职典礼引起强烈反响，导致Twitter流量猛增，如Figure 4 所示。

![](http://docs.google.com/File?id=drfcsw8_438cb3w49c9_b)

Figure 4. Twitter burst during the inauguration of Barack Obama, 1/20/2009, Tuesday   
Courtesy http://farm3.static.flickr.com/2615/4071879010\_19fb519124\_o.png

其中洪峰时刻，Twitter网站每秒钟收到350条新短信，这个流量洪峰维持了大约5分钟。根据统计，平均每个Twitter用户被120人 “追”，这就 是说，这350条短信，平均每条都要发送120次 \[16\]。这意味着，在这5分钟的洪峰时刻，Twitter网站每秒钟需要发送350 x 120 = 42,000条短信。

面对洪峰，如何才能保证网站不崩溃？办法是迅速接纳，但是推迟服务。打个比方，在晚餐高峰时段，餐馆常常客满。对于新来的顾客，餐馆服务员不是拒之门外，而是让这些顾客在休息厅等待。这就是\[14\] 所说的 “隔离用户请求与相关操作，以便烫平流量高峰”。

如何实施隔离呢？当一位用户访问Twitter网站时，接待他的是Apache Web Server。Apache做的事情非常简单，它把用户的请求解析以后，转发给Mongrel Rails Sever，由Mongrel负责实际的处理。而Apache腾出手来，迎接下一位用户。这样就避免了在洪峰期间，用户连接不上Twitter网站的尴尬局面。

虽然Apache的工作简单，但是并不意味着Apache可以接待无限多的用户。原因是Apache解析完用户请求，并且转发给 Mongrel Server以后，负责解析这个用户请求的进程(process)，并没有立刻释放，而是进入空循环，等待Mongrel Server返回结果。这样，Apache能够同时接待的用户数量，或者更准确地说，Apache能够容纳的并发的连接数量(concurrent connections)，实际上受制于Apache能够容纳的进程数量。Apache系统内部的进程机制参见Figure 5，其中每个Worker代表一个进程。

Apache能够容纳多少个并发连接呢？\[22\]的实验结果是4,000个，参见Figure 6。如何才能提高Apache的并发用户容量呢？一种思路是不让连接受制于进程。不妨把连接作为一个数据结构，存放到内存中去，释放进程，直到 Mongrel Server返回结果时，再把这个数据结构重新加载到进程上去。

事实上Yaws Web Server\[24\]，就是这么做的\[23\]。所以，Yaws能够容纳80,000以上的并发连接，这并不奇怪。但是为什么Twitter用 Apache，而不用Yaws呢？或许是因为Yaws是用Erlang语言写的，而Twitter工程师对这门新语言不熟悉 (But you need in house Erlang experience \[17\])。

![](http://docs.google.com/File?id=drfcsw8_440fdb6kkf7_b)

Figure 5. Apache web server system architecture \[21\]   
Courtesy http://farm3.static.flickr.com/2699/4071355801\_db6c8cd6c0\_o.png

![](http://docs.google.com/File?id=drfcsw8_441frq3nnd6_b)

Figure 6. Apache vs. Yaws.   
The horizonal axis shows the parallel requests,   
the vertical one shows the throughput (KBytes/second).   
The red curve is Yaws, running on NFS.   
The blue one is Apache, running on NFS,   
while the green one is also Apache but on a local file system.   
Apache dies at about 4,000 parallel sessions,   
while Yaws is still functioning at over 80,000 parallel connections. \[22\]   
Courtesy http://farm3.static.flickr.com/2709/4072077210\_3c3a507a8a\_o.jpg

Reference,

\[14\] Improving running component of Twitter.   
(http://qconlondon.com/london-2009/file?path=/qcon-london-   
2009/slides/EvanWeaver\_ImprovingRunningComponentsAtTwitter.pdf)   
\[16\] Updating Twitter without service disruptions.   
(http://gojko.net/2009/03/16/qcon-london-2009-upgrading-   
twitter-without-service-disruptions/)   
\[17\] Fixing Twitter. (http://assets.en.oreilly.com/1/event/29/Fixing\_Twitter\_   
Improving\_the\_Performance\_and\_Scalability\_of\_the\_World\_s\_Most\_Popular\_   
Micro-blogging\_Site\_Presentation%20Presentation.pdf)   
\[21\] Apache system architecture.   
(http://www.fmc-modeling.org/download/publications/   
groene\_et\_al\_2002-architecture\_recovery\_of\_apache.pdf)   
\[22\] Apache vs Yaws. (http://www.sics.se/~joe/apachevsyaws.html)   
\[23\] 质疑Apache和Yaws的性能比较. (http://www.javaeye.com/topic/107476)   
\[24\] Yaws Web Server. (http://yaws.hyber.org/)   
\[25\] Erlang Programming Language. (http://www.erlang.org/)
