AI

DeepL schools other online translators with clever machine learning

Comment

Image Credits: H. Armstrong Roberts/Getty Images

Tech giants Google, Microsoft and Facebook are all applying the lessons of machine learning to translation, but a small company called DeepL has outdone them all and raised the bar for the field. Its translation tool is just as quick as the outsized competition, but more accurate and nuanced than any we’ve tried.

I only speak a smattering of French in addition to my passable English, but luckily my colleague Frederic is a man of many tongues. We both agreed that DeepL’s translations were generally superior to those from Google Translate and Bing.

Take, for example, the following passage from a German news article, as rendered by DeepL (top) and Google:

As Frederic puts it: “Whereas Google Translate often goes for a very literal translation that misses some nuances and idioms (or gets the translation of these idioms dead wrong), DeepL often provides a more natural translation that comes closer to that of a trained translator.”

The second sentence is parsed more naturally; the measure is “designed to” accomplish something rather than just doing that thing; the police are “on the road in armoured vehicles” as opposed to merely on them; “martial appearance” may be imperfect (though inspired) but it’s far better than the nonsensical “fighters’ turmoil…had come to the fore.”

A few tests of my own on some French literature I know well enough to judge had DeepL coming out on top regularly, as well. Fewer errors of tense, intent and agreement, plus a better understanding and deployment of idiom make for a much more readable translation. We thought so, and so did translators in DeepL’s own blind testing. But don’t take anyone else’s word for it — test it out yourself.

While it’s true that meaning can be conveyed successfully despite errors of that class, as evidenced by the utility we’ve all found in even the poorest machine translations, it’s far from guaranteed that anything but the barest facts of will make it through.

Linguee evolved

DeepL was born from the similarly excellent Linguee, a translation tool that has existed for years and, while popular, never quite reached the level of Google Translate — the latter has a huge advantage in brand and position, after all. Linguee’s co-founder, Gereon Frahling, used to work for Google Research but left in 2007 to pursue this new venture.

The team has been working with machine learning for years, for tasks adjacent to the core translation, but it was only last year that they began working in earnest on a whole new system and company, both of which would bear the name DeepL.

In an email, Frahling told me that the time was ripe: “We have built a neural translation network that incorporates most of the latest developments, to which we added our own ideas.”

An enormous database of over a billion translations and queries, plus a method of ground-truthing translations by searching for similar snippets on the web, made for a strong base in the training of the new model. They also put together what they claim is the 23rd most powerful supercomputer in the world, conveniently located in Iceland.

Developments published by universities, research agencies and indeed Linguee’s competitors showed that convolutional neural networks were the way to go, rather than the recurrent neural networks the company had been using previously. Now isn’t really the place to go into the differences between CNNs and RNNs, so it must suffice to say that for accurate translation of long, complex strings of related words, the former is a better bet as long as you can control for its weaknesses.

For example, a CNN could roughly be able to be said to tackle one word of the sentence at a time. This becomes a problem when, for instance, as commonly happens, a word at the end of the sentence determines how a word at the beginning of the sentence should be formed. It’s wasteful to go through the whole sentence only to find that the first word the network picked is wrong, and then start over with that knowledge, so DeepL and others in the machine learning field apply “attention mechanisms” that monitor for such potential trip-ups and resolve them before the CNN moves on to the next word or phrase.

There are other secret techniques in play, of course, and their result is a translation tool that I personally plan to make my new default. I look forward to seeing the others step up their game.

More TechCrunch

Welcome to Startups Weekly — Haje‘s weekly recap of everything you can’t miss from the world of startups. Sign up here to get it in your inbox every Friday. Well,…

Startups Weekly: Drama at Techstars. Drama in AI. Drama everywhere.

Last year’s investor dreams of a strong 2024 IPO pipeline have faded, if not fully disappeared, as we approach the halfway point of the year. 2024 delivered four venture-backed tech…

From Plaid to Figma, here are the startups that are likely — or definitely — not having IPOs this year

Federal safety regulators have discovered nine more incidents that raise questions about the safety of Waymo’s self-driving vehicles operating in Phoenix and San Francisco.  The National Highway Traffic Safety Administration…

Feds add nine more incidents to Waymo robotaxi investigation

Terra One’s pitch deck has a few wins, but also a few misses. Here’s how to fix that.

Pitch Deck Teardown: Terra One’s $7.5M Seed deck

Chinasa T. Okolo researches AI policy and governance in the Global South.

Women in AI: Chinasa T. Okolo researches AI’s impact on the Global South

TechCrunch Disrupt takes place on October 28–30 in San Francisco. While the event is a few months away, the deadline to secure your early-bird tickets and save up to $800…

Disrupt 2024 early-bird tickets fly away next Friday

Another week, and another round of crazy cash injections and valuations emerged from the AI realm. DeepL, an AI language translation startup, raised $300 million on a $2 billion valuation;…

Big tech companies are plowing money into AI startups, which could help them dodge antitrust concerns

If raised, this new fund, the firm’s third, would be its largest to date.

Harlem Capital is raising a $150 million fund

About half a million patients have been notified so far, but the number of affected individuals is likely far higher.

US pharma giant Cencora says Americans’ health information stolen in data breach

Attention, tech enthusiasts and startup supporters! The final countdown is here: Today is the last day to cast your vote for the TechCrunch Disrupt 2024 Audience Choice program. Voting closes…

Last day to vote for TC Disrupt 2024 Audience Choice program

Featured Article

Signal’s Meredith Whittaker on the Telegram security clash and the ‘edge lords’ at OpenAI 

Among other things, Whittaker is concerned about the concentration of power in the five main social media platforms.

9 hours ago
Signal’s Meredith Whittaker on the Telegram security clash and the ‘edge lords’ at OpenAI 

Lucid Motors is laying off about 400 employees, or roughly 6% of its workforce, as part of a restructuring ahead of the launch of its first electric SUV later this…

Lucid Motors slashes 400 jobs ahead of crucial SUV launch

Google is investing nearly $350 million in Flipkart, becoming the latest high-profile name to back the Walmart-owned Indian e-commerce startup. The Android-maker will also provide Flipkart with cloud offerings as…

Google invests $350 million in Indian e-commerce giant Flipkart

A Jio Financial unit plans to purchase customer premises equipment and telecom gear worth $4.32 billion from Reliance Retail.

Jio Financial unit to buy $4.32B of telecom gear from Reliance Retail

Foursquare, the location-focused outfit that in 2020 merged with Factual, another location-focused outfit, is joining the parade of companies to make cuts to one of its biggest cost centers –…

Foursquare just laid off 105 employees

“Running with scissors is a cardio exercise that can increase your heart rate and require concentration and focus,” says Google’s new AI search feature. “Some say it can also improve…

Using memes, social media users have become red teams for half-baked AI features

The European Space Agency selected two companies on Wednesday to advance designs of a cargo spacecraft that could establish the continent’s first sovereign access to space.  The two awardees, major…

ESA prepares for the post-ISS era, selects The Exploration Company, Thales Alenia to develop cargo spacecraft

Expressable is a platform that offers one-on-one virtual sessions with speech language pathologists.

Expressable brings speech therapy into the home

The French Secretary of State for the Digital Economy as of this year, Marina Ferrari, revealed this year’s laureates during VivaTech week in Paris. According to its promoters, this fifth…

The biggest French startups in 2024 according to the French government

Spotify is notifying customers who purchased its Car Thing product that the devices will stop working after December 9, 2024. The company discontinued the device back in July 2022, but…

Spotify to shut off Car Thing for good, leading users to demand refunds

Elon Musk’s X is preparing to make “likes” private on the social network, in a change that could potentially confuse users over the difference between something they’ve favorited and something…

X should bring back stars, not hide ‘likes’

The FCC has proposed a $6 million fine for the scammer who used voice-cloning tech to impersonate President Biden in a series of illegal robocalls during a New Hampshire primary…

$6M fine for robocaller who used AI to clone Biden’s voice

Welcome back to TechCrunch Mobility — your central hub for news and insights on the future of transportation. Sign up here for free — just click TechCrunch Mobility! Is it…

Tesla lobbies for Elon and Kia taps into the GenAI hype

Crowdaa is an app that allows non-developers to easily create and release apps on the mobile store. 

App developer Crowdaa raises €1.2M and plans a US expansion

Back in 2019, Canva, the wildly successful design tool, introduced what the company was calling an enterprise product, but in reality it was more geared toward teams than fulfilling true…

Canva launches a proper enterprise product — and they mean it this time

TechCrunch Disrupt 2024 isn’t just an event for innovation; it’s a platform where your voice matters. With the Disrupt 2024 Audience Choice Program, you have the power to shape the…

2 days left to vote for Disrupt Audience Choice

The United States Department of Justice and 30 state attorneys general filed a lawsuit against Live Nation Entertainment, the parent company of Ticketmaster, for alleged monopolistic practices. Live Nation and…

Ticketmaster antitrust lawsuit could give new hope to ticketing startups

The U.K. will shortly get its own rulebook for Big Tech, after peers in the House of Lords agreed Thursday afternoon to pass the Digital Markets, Competition and Consumer bill…

‘Pro-competition’ rules for Big Tech make it through UK’s pre-election wash-up

Spotify’s addition of its AI DJ feature, which introduces personalized song selections to users, was the company’s first step into an AI future. Now, Spotify is developing an alternative version…

Spotify experiments with an AI DJ that speaks Spanish

Call Arc can help answer immediate and small questions, according to the company. 

Arc Search’s new Call Arc feature lets you ask questions by ‘making a phone call’