Searching through Unicode text with an ASCII keyboard in Python
01:51 21 Nov 2025

I have a corpus of text which includes some accented words, such as épée, and I would like people to be able to easily search through it using an ASCII keyboard. Ideally, they would simply type protege or pinata to find protégé or piñata. The program is currently written in Python and uses only the builtin libraries, such as re.

I have looked at similar questions, such as Why does re not ignore accents, but the suggested solution is to normalize the unicode string to ASCII. That could be made to work, but seems inordinately ugly and doesn't return the actual text that should be displayed. Does Python not have anything analogous to POSIX character equivalence, which maps similar characters together based on the user's locale? For example,grep -E '[[=e=]][[=p=]][[=e=]][[=e=]]' matches both epee and épée (in the en_US.UTF-8 locale).

python regex unicode ascii lc-collate