Sort Cyrillic strings before Latin in python

Viewed 1064

In my database, I have records in both Cyrillic and Latin characters. By default, they are listed alphabetically with Latin records first:

abc... bcd... cde... абв...

I would like to put the Cyrillic to the first place:

абв... abc... bcd... cde...

What I have tried so far:

  1. This solution. It is not so great because it only sorts by the first word, and I can have both Cyrillic and Latin words in the same string (or even mixed characters in the same word).

  2. Writing my own lists with Cyrillic and Latin alphabets. It works but is not great at all. I cannot take into account all possible letters in the two alphabets, including those with diacritics and write them down.

I have also been looking into PyICU but don't see how I can put it to use.

My guess is that I should use some custom collation here. The question is how this can be done in practice.

4 Answers

One way to do this is to use the transliterate module or maybe cytranslit and use a sort key that transliterates everything to the desired alphabet:

import transliterate

items = ['abc', 'bcd', 'cde', 'абв']

print(sorted(items, key=lambda x: transliterate.translit(x, 'ru')))

The output is the desired

['абв', 'abc', 'bcd', 'cde']

IMO this is not a trivial thing. I'd say that a collation is indeed required.

So, say a key function would convert a string to a tuple of codepoints, where all non-Cyrillic code points would be shifted by 100000):

import unicodedata

def key(s):
    SHIFT = 100000
    return tuple(
        ord(c) if is_cyrillic(c) else ord(c) + SHIFT
        for c in s
    )

def is_cyrillic(c):
    return unicodedata.name(c).startswith('CYRILLIC')        


>>> sorted(('wannt', 'waюnnt'), key=key)
Out[34]: ['waюnnt', 'wannt']

is_cyrillic can be optimized by using a preliminary table or caching the Cyrillic characters from the database strings.

You can try to generate a key by prepending every character with 1 if it's a Latin character and with 0 otherwise:

sorted(items, key = lambda item : ['1' + x if x < '\x7f' else '0' + x for x in item])

Very dirty but working solution for sorting according the first character. It also eliminates difference in register of letters.

ls = ['32', '24', 'xyz', 'WYZ', 'abc', 'абв', 'КЛМ', 'эюя', 'еёж', 'ёжз', '', '_']

def sort_rule(st):
    st = st.lower()
    ch = st[0] if st else ''       
    if ch >= 'а' and ch <= 'я':
        st = '1' + st
    elif ch == 'ё':
        st = '1е' + st
    elif ch >= 'a' and ch <= 'z':
        st = '2' + st
    else:
        st = '3' + st
    return st

sorted(ls, key=sort_rule)

> ['абв', 'еёж', 'ёжз', 'КЛМ', 'эюя', 'abc', 'WYZ', 'xyz', '', '24', '32', '_']

For comparison, default sorting gives the next result:

sorted(ls)

> ['', '24', '32', 'WYZ', '_', 'abc', 'xyz', 'КЛМ', 'абв', 'еёж', 'эюя', 'ёжз']
Related