APP下载

机器翻译的风险

2017-05-02ByArthurGoldhammer

英语学习 2017年4期
关键词:规则计算机人工智能

By+Arthur+Goldhammer

The ideal translator is a person “on whom nothing is lost,” said Henry James. Or maybe its a machine. But a machine wont stop you from swearing at nuns...

Years ago, on a flight from Amsterdam to Boston, two American nuns seated to my right listened to a voluble1 young Dutchman who was out to discover the United States. He asked the nuns where they were from. Alas, Framingham, Massachusetts was not on his itinerary, but, he noted, he had“shitloads of time and would be visiting shitloads of other places”.2

The jovial young Dutchman had apparently gathered that“shitloads” was a colourful synonym for the bland “lots”.3 He had mastered the syntax of English and a rather extensive vocabulary but lacked experience of the appropriateness of words to social contexts.4

This memory sprang to mind with the recent news that the Google Translate engine would move from a phrase-based system to a neural network. Both methods rely on training the machine with a “corpus”5 consisting of sentence pairs: an original and a translation. The computer then generates rules for inferring, based on the sequence6 of words in the original text, the most likely sequence of words from the target language.

The procedure is an exercise in pattern matching. Similar pattern-matching algorithms are used to interpret the syllables you utter when you ask your smartphone to “navigate to Brookline” or when a photo app tags your friends face.7 The machine doesnt “understand” faces or destinations; it reduces them to vectors8 of numbers, and processes them.

I am a professional translator, having translated some 125 books from the French. One might therefore expect me to bristle9 at Googles claim that its new translation engine is almost as good as a human translator, scoring 5.0 on a scale of 0 to 6, whereas humans average 5.1. But Im also a PhD in mathematics who has developed software that “reads” European newspapers in four languages and categorises the results by topic. So, rather than be defensive about the possibility of being replaced by a machine translator, I am aware of the remarkable feats of which machines are capable, and full of admiration for the technical complexity and virtuosity of Googles work.10

My admiration does not blind me to the shortcomings of machine translation, however. Think of the young Dutch traveler who knew “shitloads” of English. The young mans fluency demonstrated that his “wetware”—a living neural network, if you will—had been trained well enough to intuit the subtle rules (and exceptions) that make language natural.11 Computer languages, on the other hand, have context-free grammars. The young Dutchman, however, lacked the social experience with English to grasp the subtler rules that shape the native speakers diction, tone and structure. The native speaker might also choose to break those rules to achieve certain effects. If I were to say “shitloads of places”rather than “lots of places” to a pair of nuns, I would mean something by it. The Dutchman blundered into inadvertent comedy.12

Googles translation engine is “trained” on corpora ranging from news sources to Wikipedia. The bare description of each corpus is the only indication of the context from which it arises. From such scanty13 information it would be difficult to infer the appropriateness or inappropriateness of a word such as “shitloads”. If translating into French, the machine might predict a good match to beaucoup or plusieurs. This would render the meaning of the utterance but not the comedy,14 which depends on the socially marked“shitloads” in contrast to the neutral plusieurs. No matter how sophisticated the algorithm, it must rely on the information provided, and clues as to context, in particular social context, are devilishly15 hard to convey in code.

The problem, as with all previous attempts to create artificial intelligence (AI)16 going back to my student days at MIT, is that intelligence is incredibly complex. To be intelligent is not merely to be capable of inferring logically from rules or statistically from regularities. Before that, one has to know which rules are applicable, an art requiring awareness of sensitivity to situation. Programmers are very clever, but they are not yet clever enough to anticipate the vast variety of contexts from which meaning emerges. Hence even the best algorithms will miss things—and as Henry James put it, the ideal translator must be a person “on whom nothing is lost”.

This is not to say that mechanical translation is not useful. Much translation work is routine. At times, machines can do an adequate job. Dont expect miracles, however, or felicitous literary translations, or aptly rendered political zingers.17 Overconfident claims have dogged18 AI research from its earliest days. I dont say this out of fear for my job: Ive retired from translating and am devoting part of my time nowadays to…writing code.

亨利·詹姆斯說,理想的译者应该是“一无所失”之人。或者,是一无所失之机器。但是,机器可不会教你不能在修女面前爆粗口。

几年前,我从阿姆斯特丹乘机前往波士顿,两位美国修女坐在我右边,听一个正要去探索美国的荷兰小伙子侃侃而谈。他问修女从哪儿来。啊,马萨诸塞州的弗雷明汉,可惜不在他的行程计划之内。但是他说,他有“贼他妈多的时间,可以去贼他妈多的其他地方”。

这个热情友好的荷兰小伙子显然知道,“贼他妈多”跟普普通通的“很多”比起来,有趣得多。他掌握了英语的句法,有相当丰富的词汇量,却缺乏交际经验,来判断用词是否合乎语境。

想起这件事,是因为有新闻说,谷歌翻译引擎将从一个基于短语的系统,变成一个神经网络系统。两种方法都以语料库为基础,训练计算机掌握多个由原文和译文搭配组合的句子。计算机由此总结出一套规则,可以根据原句的词语排列,推导出目标语言最有可能的词语排序。

整个过程属于模式匹配的训练。当智能手机识别你的语音提问“导航到布鲁克莱恩”,或者当拍照软件识别你朋友的面部时,运用的也是类似的模式匹配算法。计算机并不能“理解”人脸或者目的地,而是把它们变成向量,再进行处理。

我是专业译者,译了差不多有125本法语书。有人因此可能会觉得,我看到谷歌的下述言论会很生气:谷歌新的翻译引擎跟人工译者一样好;若满分6分,谷歌可以打到5分,而人类的平均水平也只有5.1分。……

登录APP查看全文

猜你喜欢

规则计算机人工智能
撑竿跳规则的制定
计算机操作系统
数独的规则和演变
基于计算机自然语言处理的机器翻译技术应用与简介
人工智能与就业
让规则不规则
信息系统审计中计算机审计的应用
数读人工智能
TPP反腐败规则对我国的启示
下一幕,人工智能!