2013年1月16日星期三
搞垮银行的数学公式
经济学家费希尔·布莱克和迈伦·斯科尔斯的心血结晶——布莱克-斯科尔斯方程式(Black-Scholes equation)是投资者的圣杯,凭借这种理性方法,当金融合同正在执行时,人们就可以为其定价,如同在赛马比赛中途对一匹马的买卖下注。它开辟了一个全新的、更加复杂的投资世界,并最终孕育成为巨大的全球型产业。但是,当次级抵押贷款市场恶化时,这个金融市场的宠儿却成了“黑洞”的代名词,将世界各地的资金吸出,吞入永无止境的漩涡中。
了解危机始末的人都明白,以企业和商品为主体的实体经济被名为衍生品的复杂金融工具抢了风头。这些金融工具既不是金钱,也不是货物,而是针对投资产品的投资,在赌注上下注。衍生品构建了一个蓬勃发展的全球经济,但同时也导致市场动荡、信贷紧缩、银行体系濒临崩溃和经济衰退。正是布莱克-斯科尔斯方程式掀开了衍生品世界的潘多拉魔盒。
方程式本身并非真有问题。它很有用,也很精确,其局限性被表示得很明确,为评估金融衍生工具的可能价值提供了一个行业标准方法。因此,衍生品在到期前仍可买卖。如果明智地使用它,并在市场条件不适合的情况下弃之不用,这个公式就没问题。麻烦在于它潜在有被滥用的风险。它允许衍生品成为商品,并被人们用于谋取私利。金融界将其称为迈达斯公式(Midas Formula),视之为点物成金的秘诀。但市场忘记了迈达斯国王的故事是如何结尾的。
布莱克-斯科尔斯方程式支撑了经济的大规模增长。到2007年,在国际金融体系中的衍生品交易总量每年高达一万亿美元。扣除通胀因素,这是上世纪全球制造业总产值的十倍。该方程式的缺点是,派生出更多复杂的金融工具,其价值和风险越来越模糊不清。于是,各家公司聘请了精通数学的分析师,开发出类似的数学公式,告诉他们新工具价值几何,风险安在?然后,他们悲惨地忘记询问,一旦市场条件变化,上述答案的可靠性会有多少?
布莱克和斯科尔斯在1973年发明了该方程式,此后不久,罗伯特·默顿提供了更多的论证。它适用于最简单和最古老的衍生品:期权。期权主要有两种:看跌期权让买方有权在指定时间以商定价格出售商品。看涨期权有相似功能,但它给予的是买入而不是卖出的权利。方程式为在期权到期前计算期权价值,提供了系统化方法。因此,期权得以在任何时间买卖。方程式相当有效,以至于默顿和斯科尔斯在1997年同时获得诺贝尔经济学奖。(当时布莱克已过世,因此失去了获奖资格。)
如果每个人都知道衍生品的正确价值,而且达成共识,人们如何用它牟利?该公式要求用户对几个数值进行估算。但是,靠衍生品赚钱的主要方法是赢得赌注,也就是购买一种日后能以更高价格卖出的衍生品,或者说当衍生品到期时,其价值高于预期。赢家从输家手中赚钱利润。在任何一年,都会有75-90%的期权交易商亏钱。当次贷危机泡沫破裂时,全球的银行损失了数百亿美元。在随后的恐慌中,纳税人被迫买单,但这是政治,而不是数理经济学。
布莱克-斯科尔斯公式将期权的建议价格与四个变量相关联,其中三项可以直接代入公式,即时间、期权所含有的安全资产价值和无风险利率。这是进行零风险(如政府债券)投资时获得的纯理论收益。第四个变量是资产的波动性,以此来衡量市价变动时该期权的弹性。这个公式的假设前提是,在期权的使用期内,流动性保持不变,不需要被修正。流动性可经过价格走势的统计分析得来,但它无法以精准且万无一失的方式加以测量,而且,估算可能与现实不符。
许多金融模型背后的理念都源自生活在1900年的路易·巴舍利耶(Louis Bachelier),他认为股市波动可以为一种被称作布朗运动的随机过程。在每个瞬间,一只股票的价格要么涨,要么跌,模型假定这些事件的发生概率是固定的,也许大致相等,也许略有差异。这就像人站在街道上,反复掷硬币来决定向前或向后移动一小步,所以,他们会不停地前进后退。他们的位置也就随着股价的变化,上下起伏不定。布朗运动的最重要统计特征是它的平均值和标准差。平均值是短期均价,通常在一个特定的方向漂移,其上升或下降取决于市场对股票走势的看法。标准差可以被认为是股价偏离平均值的平均数值,可以使用标准统计公式计算出来。对于股价而言,这被称为波动率,用于衡量价格波动的不稳定性。对一段时期的股价图上,波动率与股价走势相对应。
布莱克-斯科尔斯公式将巴舍利耶的构想变为现实。它没有直接给出期权(被出售或购买时)的价值,而是用数学家所说的偏微分方程,表示出当其他多种变量发生变化时股价的变化率。幸运的是,用这个公式可以推导出看跌期权的价值计算公式,以及类似的看涨期权公式。
布莱克-斯科尔斯的早期成功,鼓励金融界发明出一系列针对不同金融工具的相关公式。传统性银行可以利用这些公式,审核贷款和交易,估算合理的利润,始终对潜在风险保持警觉。但是,非传统型银行就不那么谨慎了。很快,银行随之介入越来越多的投机冒险。
现实中的任何数学模型都依赖于简化和假设。布莱克-斯科尔斯方程式的基础是套利定价理论,其中的漂移和波动是永恒不变的。这种假设在金融理论中屡见不鲜,但对于真实市场而言,它往往是虚假的。该公式还假设交易成本为零,卖空不受限制,资金可以以已知的、固定的和无风险利率贷入贷出。现实往往再次与此大相径庭。
当这些假设有效时,风险通常很低,因为股市大波动应该是极为罕见的。但在1987年10月19日的黑色星期一,在几个小时内,全球股市价值蒸发超过20%,依照模型的假设,这种极端事件几乎是不可能发生的。金融数学专家纳西姆·尼古拉斯·塔勒布(Nassim Nicholas Taleb)在其畅销书《黑天鹅》(The Black Swan)中,将这种极端事件称为“黑天鹅”。在远古时代,所有已知的天鹅都是白色的,“黑天鹅”一词的广泛用法,就像我们现在所说的“飞天猪”。但在1697年,荷兰探险家威廉·拉明(Willem de Vlamingh)在如今被称为澳大利亚天鹅河上,发现了大群黑天鹅。因此,这个词现在用来指代一种貌似根据确凿、实际上却随时可能演变成重大错误的假设。
股市的大幅波动比布朗运动的预测更为常见。其原因是不切实际的假设 —— 忽视了潜在的黑天鹅。但在通常情况下,模型都表现得非常出色,于是,随着时间的推移和增长的信心,许多银行家和交易员忘了模型具有局限性。他们将公式当成一件护身符,保护自己在出错时免受批评。
银行、对冲基金和其他投机者很快开始了更复杂的衍生品交易,如信贷违约掉期交易,相当于为邻居的房子投保火险,而且,数目惊人。他们依据自身实力自行为资产定价,这意味着,其资产可以作为其他交易的抵押品。随着一切变得愈加复杂,用于评估价值和风险的数学模型越来越偏离现实。模型是以房地产为基础而建立的,市场认定,不动产价值将永远上升,这些投资毫无风险。
布莱克 – 斯科尔斯方程式起源于数学物理学,它认为数量无限可分,时间不断流动,变量变化平稳。这种模型对于金融世界可能是不恰当的。传统的数学经济学与现实并不总是匹配,当它出错时,会错得离谱。因此,物理学家、数学家和经济学家都在寻找更好的模型。
这些努力的最前沿成果是复杂性科学,它是数学的一个新分支,依据特定规则将市场模拟为个体互动的集合体。这些模型显示出羊群效应的破坏力:市场交易员的行为彼此效仿。上个世纪发生的每一次金融危机几乎都源于人类的羊群效应,一次又一次将世界推到崩溃的边缘,将一切毁于一旦,世界上的一座桥崩塌了,其他桥随之全部倾覆。
对生态系统的研究表明,在经济模式中,不稳定是常见的现象,其主要原因是金融体系存在设计缺陷。轻点鼠标即可完成数十亿资金的流转,可能会让利润来得更快,也可能使冲击传播得更快。
那么,应该将金融危机归咎于一个方程式吗?是,也不是。布莱克-斯科尔斯公式可能难辞其咎,但那是由于它被滥用。在任何情况下,相对于金融界的不负责任、政治的无能、不正当的激励措施和监管的疏忽而言,数学公式只在林林总总的错误中,扮演了微不足道的小角色。
尽管其所谓的专业知识,金融界的表现并不比换乱瞎猜强多少。股市已经连续20年失去方向。金融系统太过复杂,不能单凭失误频频的预言和直觉而运转,但目前的数学模型又无法充分反映现实。人们对整个系统知之甚少,而且,它岌岌可危。世界经济迫切需要一场大刀阔斧的改革,需要更多、而不是更少的数学知识。这可能是件复杂的事,但绝不是魔术。(本文摘自伊恩·斯图尔特著《改变世界的17个公式》)
美国一程序员将工作外包给中国程序员 自己每天上网冲浪
美国电信运营商Verizon的安全审计发现,美国某公司的一位顶尖程序员将自己的工作外包给中国沈阳的一家软件公司。
Verizon为该公司的雇员提供了VPN允许他们在家里工作。内部的审计发现,公司明星程序员Bob的VPN登陆日志显示他固定的从中国沈阳访问公司主服务器。调查排除了恶意程序或黑客入侵的可能性。Verizon的进一步分析发现,Bob雇用了沈阳的一家软件咨询公司去做他的日常编程工作,通过FedExed提供了他的双步认证令牌,将薪水的五分之一支付给外包公司,自己则去上网冲浪,比如在Reddit上逛几小时,然后去吃饭,再上ebay血拼,更新Facebook 和LinkedIn,最后给经理发送每日工作邮件,上床睡觉。他的计划一直行之有效,在公司人力资源眼里他是最高效的程序员之一,被认为是C, C++、Perl、Java、Ruby, PHP和Python方面的专家。目前Bob已被解雇。
———
Security audit finds dev OUTSOURCED his JOB to China
Cunning scheme netted him ‘best in company’ awards
By Iain Thomson in San Francisco • Get more from this author
Posted in Security, 16th January 2013 01:29 GMT
A security audit of a US critical infrastructure company last year revealed that its star developer had outsourced his own job to a Chinese subcontractor and was spending all his work time playing around on the internet.
The firm’s telecommunications supplier Verizon was called in after the company set up a basic VPN system with two-factor authentication so staff could work at home. The VPN traffic logs showed a regular series of logins to the company’s main server from Shenyang, China, using the credentials of the firm’s top programmer, “Bob”.
“The company’s IT personnel were sure that the issue had to do with some kind of zero day malware that was able to initiate VPN connections from Bob’s desktop workstation via external proxy and then route that VPN traffic to China, only to be routed back to their concentrator,” said Verizon. “Yes, it is a bit of a convoluted theory, and like most convoluted theories, an incorrect one.”
After getting permission to study Bob’s computer habits, Verizon investigators found that he had hired a software consultancy in Shenyang to do his programming work for him, and had FedExed them his two-factor authentication token so they could log into his account. He was paying them a fifth of his six-figure salary to do the work and spent the rest of his time on other activities.
The analysis of his workstation found hundreds of PDF invoices from the Chinese contractors and determined that Bob’s typical work day consisted of:
9:00 a.m. – Arrive and surf Reddit for a couple of hours. Watch cat videos
11:30 a.m. – Take lunch
1:00 p.m. – Ebay time
2:00-ish p.m – Facebook updates, LinkedIn
4:30 p.m. – End-of-day update e-mail to management
5:00 p.m. – Go home
The scheme worked very well for Bob. In his performance assessments by the firm’s human resources department, he was the firm’s top coder for many quarters and was considered expert in C, C++, Perl, Java, Ruby, PHP, and Python.
Further investigation found that the enterprising Bob had actually taken jobs with other firms and had outsourced that work too, netting him hundreds of thousands of dollars in profit as well as lots of time to hang around on internet messaging boards and checking out the latest Detective Mittens video.
Bob is no longer employed by the firm. ®
长尾理论
长尾理论
求助编辑百科名片
长尾理论(The Long Tail)是网络时代兴起的一种新理论,由美国人克里斯·安德森提出。长尾理论认为,由于成本和效率的因素,当商品储存流通展示的场地和渠道足够宽广,商品生产成本急剧下降以至于个人都可以进行生产,并且商品的销售成本急剧降低时,几乎任何以前看似需求极低的产品,只要有卖,都会有人买。这些需求和销量不高的产品所占据的共同市场份额,可以和主流产品的市场份额 相比,甚至更大。
编辑本段由来及含义
根据维基百科,长尾(The Long Tail)这一概念是由《连线》杂志主编Chris Anderson在2004年十月的“长尾” 一文中最早提出,用来描述诸如亚马逊和Netflix之类网站的商业和经济模式。
过去人们只能关注重要的人或重要的事,如果用正态分布曲线来描绘这些人或事,人们只能关注曲线的“头部”,而将处于曲线“尾部”、需要更多的精力和成本才能关注到的大多数人或事忽略。例如,在销售产品时,厂商关注的是少数几个所谓“VIP”客户,“无暇”顾及在人数上居于大多数的普通消费者。而在网络时代,由于关注的成本大大降低,人们有可能以很低的成本关注正态分布曲线的“尾部”,关注“尾部”产生的总体效益甚至会超过“头部”。例如,某著名网站是世界上最大的网络广告商,它没有一个大客户,收入完全来自被其他广告商忽略的中小企业。安德森认为,网络时代是关注“长尾”、发挥“长尾”效益的时代。
举例来说,我们常用的汉字实际上不多,但因出现频次高,所以这些为数不多的汉字占据了上图广大
的红区;绝大部分的汉字难得一用,它们就属于长尾。 Chris认为,只要存储和流通的渠道足够大,需求不旺或销量不佳的产品共同占据的市场份额就可以和那些数量不多的热卖品所占据的市场份额相匹敌甚至更大。
长尾市场也称之为“利基市场”。“利基”一词是英文“Niche” 的音译,意译为“壁龛”,有拾遗补缺或见缝插针的意思。菲利普·科特勒在《营销管理》中给利基下的定义为:利基是更窄地确定某些群体,这是一个小市场并且它的需要没有被服务好,或者说“有获取利益的基础”。
通过对市场的细分,企业集中力量于某个特定的目标市场,或严格针对一个细分市场,或重点经营一个产品和服务,创造出产品和服务优势。
编辑本段内涵
简单的说,所谓长尾理论是指,只要产品的存储和流通的渠道足够大,需求不旺或销量不佳的产品所共同占据的市场份额可以和那些少数热销产品所占据的市场份额相匹敌甚至更大,即众多小市场汇聚成可产生与主流相匹敌的市场能量。也就是说,企业的销售量不在于传统需求曲线上那个代表“畅销商品”的头部,而是那条代表“冷门商品”经常为人遗忘的长尾。举例来说,一家大型书店通常可摆放10万本书,但亚马逊网络书店的图书销售额中,有四分之一来自排名10万以后的书籍。这些“冷门”书籍的销售比例正以高速成长,预估未来可占整个书市的一半。这意味着消费者在面对无限的选择时,真正想要的东西、和想要取得的渠道都出现了重大的变化,一套崭新的商业模式也跟着崛起。简而言之,长尾所涉及的冷门产品涵盖了几乎更多人的需求,当有了需求后,会有更多的人意识到这种需求,从而使冷门不再冷门。
编辑本段发现
克里斯·安德森,美国《连线》杂志主编,喜欢从数字中发现趋势。一次跟eCast首席执行官范·阿迪布的会面,后者提出一个让安德森耳目一新的“98法则”,改变了他的研究方向。范·阿迪布从数字音乐点唱数字统计中发现了一个秘密:听众对98的非热门音乐有着无限的需求,非热门的音乐集合市场无比巨大,无边无际。听众几乎盯着所有的东西!他把这称为“98法则”。
安德森意识到阿迪布那个有悖常识的“98法则”,隐含着一个强大的真理。于是,他系统研究了亚马逊、狂想曲公司、Blog、Google、eBay、Netflix等互联网零售商的销售数据,并与沃尔玛等传统零售商的销售数据进行了对比,观察到一种符合统计规律(大数定律)的现象。这种现象恰如以数量、品种二维坐标上的一条需求曲线,拖着长长的尾巴,向代表“品种”的横轴尽头延伸,长尾由此得名。
《长尾》(long tail)在2004年10月号《连线》发表后,迅速成了这家杂志历史上被引用最多的一篇文章。特别是经过吸纳无边界智慧的博客平台,不断丰富着新的素材和案例。安德森沉浸其中不能自拔,终于打造出一本影响商业世界的畅销书《长尾理论》。
编辑本段案例
Google adwords、Amazon、Itune都是长尾理论的优秀案例。
1、 Google是一个最典型的“长尾”公司,其成长历程就是把广告商和出版商的“长尾”商业化的过
程。以占据了Google半壁江山的AdSense为例,它面向的客户是数以百万计的中小型网站和个人—对于普通的媒体和广告商而言,这个群体的价值微小得简直不值一提,但是Google通过为其提供个性化定制的广告服务,将这些数量众多的群体汇集起来,形成了非常可观的经济利润。目前,Google的市值已超过2100亿美元,被认为是“最有价值的媒体公司”,远远超过了那些传统的老牌传媒。
2、长尾理论与图书出版。图书出版业是“小众产品”行业,市场上流通的图书达300万种。大多数图书很难找到自己的目标读者,只有极少数的图书最终成为畅销书。由于长尾书的印数及销量少,而出版、印刷、销售及库存成本又较高,因此,长期以来出版商和书店的经营模式多以畅销书为中心。网络书店和数字出版社的发展为长尾书销售提供了无限的空间市场。在这个市场里,长尾书的库存和销售成本几乎为零,于是,长尾图书开始有价值了。销售成千上万的小众图书,哪怕一次仅卖一两本,其利润累计起来可以相当甚至超过那些动辄销售几百万册的畅销书。如亚马逊副经理史蒂夫·凯塞尔所说:“如果我有10万种书,哪怕一次仅卖掉一本,10年后加起来它们的销售就会超过最新出版的《哈利·波特》。”
编辑本段与二八定律
“长尾理论”被认为是对传统的“二八定律”的彻底叛逆。
尽管听上去有些学术的味道,但事实上这不难理解——人类一直在用二八定律来界定主流,计算投入和产出的效率。它贯穿了整个生活和商业社会。这是1897年意大利经济学家帕累托归纳出的一个统计结论,即20%的人口享有80%的财富。当然,这并不是一个准确的比例数字,但表现了一种不平衡关系,即少数主流的人(或事物)可以造成主要的、重大的影响。以至于在市场营销中,为了提高效率,厂商们习惯于把精力放在那些有80%客户去购买的20%的主流商品上,着力维护购买其20%商品的80%的主流客户。
在上述理论中被忽略不计的80%就是长尾。Chris Anderson说:“我们一直在忍受这些最小公分母的专制统治……我们的思维被阻塞在由主流需求驱动的经济模式下。”但是人们看到,在互联网的促力下,被奉为传统商业圣经的“二八定律”开始有了被改变的可能性。这一点在媒体和娱乐业尤为明显,经济驱动模式呈现从主流市场向非主流市场转变的趋势。
长尾理论无处不在?长尾理论的应用决不止于互联网以及娱乐媒体产业。
传统的市场曲线是符合80/20铁律的,为了抢夺那带来80% 利润的畅销品市场,我们厮杀得天昏地暗,但是我们所谓的热门商品正越来越名不副实,比如说黄金电视节目的收视率几十年来一直在萎缩,若放在1970年,现在的一档最佳节目恐怕连前10名之列都难以进入。简言之,尽管我们仍然对大热门着迷,但它们的经济力量已经今非昔比。那么,那些反复无常的消费者们已经转向了什么地方?答案并非唯一。他们散向了四面八方,因为市场已经分化成了无数不同的领域。互联网的出现改变了这种局面,使得99%的商品都有机会进行销售,市场曲线中那条长长的尾部(所谓的利基产品)也咸鱼翻身,成为我们可以寄予厚望的新的利润增长点。
请看这张统计,横轴是品种,纵轴是销量。典型的情况是只有少数产品销量较高,其余多数产品销量很低。传统的二八定律(或称20/80定律)关注其中红色部分,认为20%的品种带来了80%的销量,所以应该只保留这部分,其余的都应舍弃。长尾理论则关注蓝色的长尾巴,认为这部分积少成多,可以积累成足够大、甚至超过红色部分的市场份额。但也有很多失败者并没有真正理解长尾理论的实现条件。
首先,长尾理论统计的是销量,并非利润。管理成本是其中最关键的因素。销售每件产品需要一定的成本,增加品种所带来的成本也要分摊。所以,每个品种的利润与销量成正比,当销量低到一个限度就会亏损。理智的零售商是不会销售引起亏损的商品。这就是二八定律的基础。
超市是通过降低单品销售成本,从而降低每个品种的止亏销量,扩大销售品种。为了吸引顾客和营造货品齐全的形象,超市甚至可以承受亏损销售一些商品。但迫于仓储、配送的成本,超市的承受能力是有限的。
互联网企业可以进一步降低单品销售成本,甚至没有真正的库存,而网站流量和维护费用远比传统店面低,所以能够极大地扩大销售品种。比如Amazon就是如此。而且,互联网经济有赢者独占的特点,所以网站在前期可以不计成本、疯狂投入,这更加剧了品种的扩张。
如果互联网企业销售的是虚拟产品,则支付和配送成本几乎为0,可以把长尾理论发挥到极致。Google adwords、iTunes音乐下载都属于这种情况。可以说,虚拟产品销售天生就适合长尾理论。
其次,要使长尾理论更有效,应该尽量增大尾巴。也就是降低门槛,制造小额消费者。不同于传统商业的拿大单、传统互联网企业的会员费,互联网营销应该把注意力放在把蛋糕做大。通过鼓励用户尝试,将众多可以忽略不计的零散流量,汇集成巨大的商业价值。
Google adsense就是这样一个蛋糕制造机。之前,普通个人网站几乎没有盈利机会。Adsense通过在小网站上发布相关广告,带给站长们一种全新的低门槛的盈利渠道。同时,把众多小网站的流量汇集成为统一的广告媒体。
当然,在这里还有一个降低管理成本的问题。如果处理不好,客服成本会迅速上升,成为主要矛盾。Google是通过算法降低人工管理工作量,但也仅仅做到差强人意。
使用长尾理论必须小心翼翼,保证任何一项成本都不随销量的增加而激增,最差也是同比增长。否则,就会走入死路。最理想的长尾商业模式是,成本是定值,而销量可以无限增长。这就需要可以低成本扩展的基础设施,Google的bigTable就是如此。
编辑本段长尾理论给新媒体带来的影响
第一,互联网为新媒体传播提供了无限的空间市场,任何曾经创造的内容原则上都将在这里"永生"。传统媒体在传播过程中,从经济上考虑,销售商不可能去经营太多的处于长尾的小众商品。互联网平台为这些长尾小众商品提供了销售市场。所有非主流的市场累加起来就会形成一个比主流市场还大的市场,这些少量的需求会在需求曲线上面形成一条长长的“尾巴”,实现小众的极大数量。在长长的“尾巴"上,曾被大众流行挤压和忽略的"个性化"将被凸现出来。
第二,从制作和传播上来说,传统媒体的制作和传播成本是相对高昂的。在互联网上,网民可以不花分文上传网页或撰写博客,还可以免费在网络上传播自己的内容。低成本的制作和传播将会使从事长尾小众商品的生产和传播者获得更好的利益回报,从而繁荣长尾小众商品的供应市场。
第三,传统媒体是一种内容打包的服务,一方面众口难调,所以要用各种文章的集合最大化地满足最多的读者,另一方面,内容的传播速度较慢,修改和更新也不方便,读者很难实时地获得某些信息。而互联网时代,读者可以随时用他感兴趣的关键词搜索,看他想看的文章,甚至可以实时获得某些重要信息,而这些文章和信息就可能来自于处于长尾的小众商品。利用RSS技术,人们还可以在互联网上打造一份自己的个性化报纸。于是,互联网就可能用每一篇文章来满足每一个读者的需求。[1]
编辑本段长尾理论认为的传播力量来源
1、智能终端(智能手机、个人电脑、平板电脑等)的普及使内容生产普及,廉价的生产得以实现
2、互联网传播工具的普及,消费和营销成本显著下降
3、搜索引擎把低成本的产品和少量可能的无限需求迅速连接起来,使需求曲线向尾部移动
编辑本段长尾理论及其对搜索引擎营销策略的意义
从上述示意图中可以看出,与20/80定律不同是,长尾理论中“尾巴”的作用是不能忽视的,经营者不应该只关注头部的作用。长尾理论已经成为一种新型的经济模式,被成功应用于网络经济领域。举例来说,Google就有效地利用了长尾策略。google的Adwords广告使得无数中小企业都能自如投放网络广告,而传统的网络广告投放只是大企业才能涉足的领域。其Adsense广告又使得大批中小网站都能自动获得广告商投放广告。Adwords和Adsense因此汇聚成千上万的中小企业和中小网站,其产生的巨大价值和市场能量足以抗衡传统网络广告市场。如果google只是将市场的注意力放在20%的大企业身上(像许多门户网站的网络广告策略那样),那么也很难创造现在的辉煌了。同样,网上零售巨人亚马逊的商品包罗万象,而不仅仅是那些可以创造高利润的少数商品,结果证明,亚马逊模式是成功的,而那些忽视长尾,仅仅关注少数畅销商品的网站经营状况并不理想。
长尾理论对于搜索引擎营销中的关键词策略非常有用。即虽然少数核心关键词或通用关键词可以为网站带来可能超过一半的访问量,但那些搜索人数不多然而非常明确的关键词的总和——即长尾关键词同样能为网站带来可观的访问量,并且这些长尾关键词检索所形成的顾客转化率更高,往往也大大高于通用关键词的转化率。比如,一个利用通用词汇“律师”进行检索到达网站访问者与一个搜索“北京商标权纠纷律师”到达网站的访问者相比,后者更加容易转化成该网站的客户。这也就是研究用户关键词检索行为分散性以及分散关键词策略的价值所在。
2013年1月15日星期二
英媒:为什么一半中国孕妇剖腹产?
- 《每日电讯报》网络版周一(14日)刊登署名Tessa Thorniley的文章,谈为什么一半中国女性选择剖腹产。
- 作者说,她的采访对象是一位讲一口流利中文的美国人MK,因为在上海的一家顶级私人医院做助产士,MK对中国产妇的心态、观点和生活都很了解。
- MK接触的都是中国最有钱的孕妇,她的第一印象就是,中国文化和西方文化在分娩问题上存在巨大差异。
- 传统上说,中国女性一般在任何方面都能忍受自己个人的不适,包括疼痛,但一旦面临分娩,她们的忍耐力就似乎消失了。
- 剖腹产
- 据一项对中国产科医院的统计,在2007到2008年间,大约有50%的中国产妇选择剖腹产,为全世界剖腹产比率最高。
- 即使那些自然分娩的产妇,比如城市产妇和能到医院生孩子的农村产妇,也会坚持要求使用麻醉药分娩。
- MK说,出现这种情况的首要原因是作选择的往往不是产妇本人,而是产科医院,因为中国的公立医院实在太拥挤,一天可做10到12个剖腹产,却只能做2到3个自然分娩。
- 在医院床位有限的同时,中国各地的医院还大量缺乏助产士,所以自然分娩的接生也会有问题。
- 产科医生喜欢剖腹产的另一个原因是避免产房外面过于拥挤。每个产妇几乎都有一大家子人在产房外等着新生儿出生,剖腹产的产妇推进去,只需30分钟就能解决问题,家人就会离开。
- 但想象一下,如果所有产妇都自然分娩,那么产房就会被挤破了门,没有人能赶走那些焦急等待的家人。
- 此外,由于中国目前仍然推行独生子女政策,而一般人都认为剖腹产比自然分娩安全,所以为了让他们唯一的孩子能安全出生,人们通常也会选择剖腹产。
- 还有一个原因是迷信,一些特定的年月,被认为是孩子出生的好日子,而另一些日子则被认为在那时出生的孩子一辈子会倒霉,剖腹产则能选择好日子分娩,避免孩子将来有厄运。
- 麻醉药
- MK还透露,她接触的产妇中,有越来越多人为了害怕分娩时的痛苦而选择剖腹产,不剖腹的也坚持使用麻醉药。
- 这些产妇都是在一胎化政策施行后出生的独生女,他们在父母和祖父母的溺爱下长大,完全不同于中国妇女对自身不适有很强忍耐力的传统。
- 但在那些非常贫穷地区的孕妇们,她们通常只能在家生孩子,没有医生也没有助产士,只能忍受所有的疼痛,有缺陷新生儿的比率也非常大。
- 不过MK指出,其实在发达国家的女性中,对待分娩疼痛问题也有很多不同的观点和做法。
- 在西欧,荷兰产妇在家生孩子的比率非常高,达65%,她们不使用麻醉药,但法国产妇比较喜欢麻醉药,英国的则是对半。
2013年1月14日星期一
New Challenges in Computer Science Research
Posted by Jeff Walz, Head of University Relations
Yesterday afternoon at the 2012 Computer Science Faculty Summit, there was a round of lightning talks addressing some of the research problems faced by Google across several domains. The talks pointed out some of the biggest challenges emerging from increasing digital interaction, which is this year’s Faculty Summit theme.
Research Scientist Vivek Kwatra kicked things off with a talk about video stabilization on YouTube. The popularity of mobile devices with cameras has led to an explosion in the amount of video people capture, which can often be shaky. Vivek and his team have found algorithmic approaches to make casual videos look more professional by simulating professional camera moves. Their stabilization technology vastly improves the quality of amateur footage.
Next, Ed Chi (Research Scientist) talked about social media focusing on the experimental circle model that characterizes Google+. Ed is particularly interested in how social interaction on the web can be designed to mimic live communication. Circles on Google+ allow a user to manage their audience and share content in a targeted fashion, which reflects face-to-face interaction. Ed discussed how, from an HCI perspective, the challenge going forward is the need to consider the trinity of social media: context, audience, content.
John Wilkes, Principal Software Engineer, talked about cluster management at Google and the challenges of building a new cluster manager-- that is, an operating system for a fleet of machines. Everything at Google is big and a consequence of operating at such tremendous scale is that machines are bound to fail. John’s team is working to make things easier for internal users enabling our ability to respond to more system requests. There are several hard problems in this domain, such as issues with configuration, making it as easy as possible to run a binary, increasing failure tolerance, and helping internal users understand their own needs as well as the behavior and performance of their system in our complicated distributed environment.
Research Scientist and coffee connoisseur Alon Halevy took to the podium to confirm that he did indeed author an empirical book on coffee, and also talked with attendees about structured data on the web. Structured data is comprised of hundreds of millions of (relatively small) tables of data, and Alon’s work is focused on enabling data enthusiasts to discover and visualize those data sets. Great possibilities open up when people start combining data sets in meaningful ways, which inspired the creation of Fusion Tables. An example is a map made in the aftermath of the 2011 earthquake and tsunami in Japan, that shows natural disaster data alongside the locations of the world’s nuclear plants. Moving forward, Alon’s team will continue to think about interesting things that can be done with data, and the techniques needed to distinguish good data from bad data.
To wrap up the session, Praveen Paritosh did a brief, but deep dive into the Knowledge Graph, an intelligent model that understands real-world entities and their relationships to one another-- things, not strings-- which launched earlier this year.
The Google Faculty Summit continued today with more talks, and breakout sessions centered on our theme of digital interaction. Check back for additional blog posts in the coming days.
Yesterday afternoon at the 2012 Computer Science Faculty Summit, there was a round of lightning talks addressing some of the research problems faced by Google across several domains. The talks pointed out some of the biggest challenges emerging from increasing digital interaction, which is this year’s Faculty Summit theme.
Research Scientist Vivek Kwatra kicked things off with a talk about video stabilization on YouTube. The popularity of mobile devices with cameras has led to an explosion in the amount of video people capture, which can often be shaky. Vivek and his team have found algorithmic approaches to make casual videos look more professional by simulating professional camera moves. Their stabilization technology vastly improves the quality of amateur footage.
Next, Ed Chi (Research Scientist) talked about social media focusing on the experimental circle model that characterizes Google+. Ed is particularly interested in how social interaction on the web can be designed to mimic live communication. Circles on Google+ allow a user to manage their audience and share content in a targeted fashion, which reflects face-to-face interaction. Ed discussed how, from an HCI perspective, the challenge going forward is the need to consider the trinity of social media: context, audience, content.
John Wilkes, Principal Software Engineer, talked about cluster management at Google and the challenges of building a new cluster manager-- that is, an operating system for a fleet of machines. Everything at Google is big and a consequence of operating at such tremendous scale is that machines are bound to fail. John’s team is working to make things easier for internal users enabling our ability to respond to more system requests. There are several hard problems in this domain, such as issues with configuration, making it as easy as possible to run a binary, increasing failure tolerance, and helping internal users understand their own needs as well as the behavior and performance of their system in our complicated distributed environment.
Research Scientist and coffee connoisseur Alon Halevy took to the podium to confirm that he did indeed author an empirical book on coffee, and also talked with attendees about structured data on the web. Structured data is comprised of hundreds of millions of (relatively small) tables of data, and Alon’s work is focused on enabling data enthusiasts to discover and visualize those data sets. Great possibilities open up when people start combining data sets in meaningful ways, which inspired the creation of Fusion Tables. An example is a map made in the aftermath of the 2011 earthquake and tsunami in Japan, that shows natural disaster data alongside the locations of the world’s nuclear plants. Moving forward, Alon’s team will continue to think about interesting things that can be done with data, and the techniques needed to distinguish good data from bad data.
To wrap up the session, Praveen Paritosh did a brief, but deep dive into the Knowledge Graph, an intelligent model that understands real-world entities and their relationships to one another-- things, not strings-- which launched earlier this year.
The Google Faculty Summit continued today with more talks, and breakout sessions centered on our theme of digital interaction. Check back for additional blog posts in the coming days.
Machine Learning Book for Students and Researchers
Posted by Afshin Rostamizadeh, Google Research
Our machine learning book, The Foundations of Machine Learning, is now published! The book, with authors from both Google Research and academia, covers a large variety of fundamental machine learning topics in depth, including the theoretical basis of many learning algorithms and key aspects of their applications. The material presented takes its origin in a machine learning graduate course, "Foundations of Machine Learning", taught by Mehryar Mohri over the past seven years and has considerably benefited from comments and suggestions from students and colleagues at Google.
The book can serve as a textbook for both graduate students and advanced undergraduate students and a reference manual for researchers in machine learning, statistics, and many other related areas. It includes as a supplement introductory material to topics such as linear algebra and optimization and other useful conceptual tools, as well as a large number of exercises at the end of each chapter whose full solutions are provided online.
Our machine learning book, The Foundations of Machine Learning, is now published! The book, with authors from both Google Research and academia, covers a large variety of fundamental machine learning topics in depth, including the theoretical basis of many learning algorithms and key aspects of their applications. The material presented takes its origin in a machine learning graduate course, "Foundations of Machine Learning", taught by Mehryar Mohri over the past seven years and has considerably benefited from comments and suggestions from students and colleagues at Google.
The book can serve as a textbook for both graduate students and advanced undergraduate students and a reference manual for researchers in machine learning, statistics, and many other related areas. It includes as a supplement introductory material to topics such as linear algebra and optimization and other useful conceptual tools, as well as a large number of exercises at the end of each chapter whose full solutions are provided online.
Better table search through Machine Learning and Knowledge
Posted By Johnny Chen, Product Manager, Google Research
The Web offers a trove of structured data in the form of tables. Organizing this collection of information and helping users find the most useful tables is a key mission of Table Search from Google Research. While we are still a long way away from the perfect table search, we made a few steps forward recently by revamping how we determine which tables are "good" (one that contains meaningful structured data) and which ones are "bad" (for example, a table that hold the layout of a Web page). In particular, we switched from a rule-based system to a machine learning classifier that can tease out subtleties from the table features and enables rapid quality improvement iterations. This new classifier is asupport vector machine (SVM) that makes use of multiple kernel functions which are automatically combined and optimized using training examples. Several of these kernel combining techniques were in fact studied and developed within Google Research [1,2].
We are also able to achieve a better understanding of the tables by leveraging the Knowledge Graph. In particular, we improved our algorithms for identifying the context and topics of each table, the entities represented in the table and the properties they have. This knowledge not only helps our classifier make a better decision on the quality of the table, but also enables better matching of the table to the user query.
Finally, you will notice that we added an easy way for our users to import Web tables found through Table Search into their Google Drive account as Fusion Tables. Now that we can better identify good tables, the import feature enables our users to further explore the data. Once in Fusion Tables, the data can be visualized, updated, and accessed programmatically using the Fusion Tables API.
These enhancements are just the start. We are continually updating the quality of our Table Search and adding features to it.
Stay tuned for more from Boulos Harb, Afshin Rostamizadeh, Fei Wu, Cong Yu and the rest of the Structured Data Team.
[1] Algorithms for Learning Kernels Based on Centered Alignment
[2] Generalization Bounds for Learning Kernels
The Web offers a trove of structured data in the form of tables. Organizing this collection of information and helping users find the most useful tables is a key mission of Table Search from Google Research. While we are still a long way away from the perfect table search, we made a few steps forward recently by revamping how we determine which tables are "good" (one that contains meaningful structured data) and which ones are "bad" (for example, a table that hold the layout of a Web page). In particular, we switched from a rule-based system to a machine learning classifier that can tease out subtleties from the table features and enables rapid quality improvement iterations. This new classifier is asupport vector machine (SVM) that makes use of multiple kernel functions which are automatically combined and optimized using training examples. Several of these kernel combining techniques were in fact studied and developed within Google Research [1,2].
We are also able to achieve a better understanding of the tables by leveraging the Knowledge Graph. In particular, we improved our algorithms for identifying the context and topics of each table, the entities represented in the table and the properties they have. This knowledge not only helps our classifier make a better decision on the quality of the table, but also enables better matching of the table to the user query.
Finally, you will notice that we added an easy way for our users to import Web tables found through Table Search into their Google Drive account as Fusion Tables. Now that we can better identify good tables, the import feature enables our users to further explore the data. Once in Fusion Tables, the data can be visualized, updated, and accessed programmatically using the Fusion Tables API.
These enhancements are just the start. We are continually updating the quality of our Table Search and adding features to it.
Stay tuned for more from Boulos Harb, Afshin Rostamizadeh, Fei Wu, Cong Yu and the rest of the Structured Data Team.
[1] Algorithms for Learning Kernels Based on Centered Alignment
[2] Generalization Bounds for Learning Kernels
订阅:
博文 (Atom)

