Yandex’s Algorithm Chases Its Own Tail
At a recent talk at International Search Summit, a Yandex representative summed up their basic algorithm: for a given query or category, they’ll pick an “archetype” site (or sites), and then they use machine learning to rank other sites based on their similarity to the first site.
In some ways, this elegantly disposes of gaming: let’s say the archetype is always careful to use the “keywords” attribute to list terms relevant to each page. Other sites will copy that, and careful usage of the “keywords” attribute will end up being statistically meaningless. Perfect: harmful (or at least pointless) gaming gets disposed of.
So the search engine gamers move on to the next signal. Let’s say they notice that the archetype sites are all well laid-out, and present useful information in a visually enticing format. So they copy that, and—whoops: the higher site quality goes, the weaker a signal it gets—notwhat you’d want from a search engine.
This kind of algorithm probably has plenty of safeguards against this kind of risk. And by the time good layouts have zero effect on rankings, they’re so ubiquitous that it doesn’t matter. But search is a very adaptive market; the interesting stuff happens on the margin. And on the margin, the highest risk/reward activity for gaming this kind of algorithm is to guess a new search factor, rather than optimizing an old search factor.
Lots of search engines—Yandex and Google included—can get by with imperfect algorithms for the moment. The Internet is growing fast, and lots of people build good sites for their own sake. But gamable search engines are always at risk from even marginally better competitors.
And that’s not even the big risk Yandex faces: the most rewarding way to game Yandex-style algorithms is to try to become canonical, whatever that might mean.
And that’s the real weakness, here: in the end, these search engines are “seeded” with data from a different search algorithm—a sort of pagerank for prestige, with all the benefits and drawbacks that entails. This gives these search engines a huge head start, but means that they’re fundamentally limited in how much they can improve; historically, better data will beat better algorithms (and winners can hire good algorithm-writers after the fact to grow their lead).
Being the #1 search engine in a fast-growing market is a major accomplishment. Yandex deserves credit for taking on bigger competitors, without having to copy them. But they’re in a tough situation: Google started by ignoring the kind of data Yandex is built on. Now, they’re incorporating more and more of it. To become a more brand-focused search engine is a simple, incremental decision for Google; to become Googlier is a fundamental change for Yandex.
Yandex’s management team is too smart to bet against, and their record is undeniable. But every stable stream of free cash flow is a bet on the status quo, and for Yandex, that status quo is more tenuous than it looks.
Full disclosure: I work for Yahoo, so my employer competes with Yandex in an indirect and insignificant way. We’re continuing to experiment with new formats for Digital Due Diligence. A weekly link roundup and more-than-weekly editorials seems ideal thus far. As always, we welcome your comments.
Recent Research
Digital Due Diligence Weekly
