IndicPhone

IndicPhone is a universal algorithm for phonetically hashing Indic-language words written in all Brahmic scripts in the Unicode table, similar to the Metaphone algorithm for English. For a given word, it generates three Romanized phonetic keys (hashes), each representing an increasing degree of affinity to various phonetic properties of the original word. It was originally written in 2013 for enabling fuzzy search in the Olam Malayalam dictionary.

The algorithm generates phonetic hashes for Indic words by accounting for linguistic features such as compounding and gemination. It exploits the consistent phonetic ordering of glyphs (derived from the ISCII-based Unicode layout) in Brahmic scripts in their respective Unicode blocks, to transform all Indic glyphs in a single pass using a universal mapping of phonetic offsets to alphanumeric Roman characters.

Demo

Examples

Input key0 key1 key2
Kamala
कमलKMLKMLKML
কমলKMLKMLKML
ಕಮಲKMLKMLKML
കമലKMLKMLKML
Svatantra
स्वतन्त्रSV0N0RSV0N0RSV0N0R
স্বতন্ত্রSB0N0RSB0N0RSB0N0R
ಸ್ವತನ್ತ್ರSV0N0RSV0N0RSV0N0R
സ്വതന്ത്രSV0N0RSV0N0RSV0N0R
Namaste
नमस्तेNMS0NMS0NMS06
নমস্তেNMS0NMS0NMS06
ನಮಸ್ತೇNMS0NMS0NMS06
നമസ്തേNMS0NMS0NMS06
Nīlakkuyil
नीलक्कुयिल्NLKYLNLKYLN4LK25Y4L
নীলক্কুযিল্NLKYLNLKYLN4LK25Y4L
ನೀಲಕ್ಕುಯಿಲ್NLKYLNLKYLN4LK25Y4L
നീലക്കുയിൽNLKYLNLKYLN4LK25Y4L
Gaṅgā
गंगK3KK3KK3K
গংগK3KK3KK3K
ಗಂಗK3KK3KK3K
ഗംഗK3KK3KK3K
Himālaya
हिमालयHMLYHMLYH4MLY
হিমালয়HMLYHMLYH4MLY
ಹಿಮಾಲಯHMLYHMLYH4MLY
ഹിമാലയHMLYHMLYH4MLY

The algorithm

IndicPhone converts a word in any Indic/Brahmic script into three phonetic hashes in a single left-to-right pass. The crux of the algorithm is a universal phonetic mapping table that is shared across every script, which is used to map sounds to specific Roman characters.

1. Universal translation table

Each Brahmic script has its own distinctive 128-point Unicode block, where all blocks follow the same ISCII-derived phonetic ordering. A given offset (the position of a glyph within a block) represents the same sound in every script. For example, offset 0x15 is क in Devanagari, ক in Bengali, ಕ in Kannada, and ക in Malayalam, all the same K sound.

offsetDevanagariBengaliKannadaMalayalamCode
Vowels
0x5A
0x6A
0x7I
0x8I
0x9U
0xaU
0xbR
0xeE
0xfE
0x10AI
0x12O
0x13O
0x14O
0x60R
Consonants
0x15K
0x16K
0x17K
0x18K
0x19NG
0x1aC
0x1bC
0x1cJ
0x1dJ
0x1eNJ
0x1fT
0x20T
0x21T
0x22T
0x23N1
0x240
0x250
0x260
0x270
0x28N
0x2aP
0x2bF
0x2cB
0x2dB
0x2eM
0x2fY
0x30R
0x31R1
0x32L
0x33L1
0x34Z
0x35V
0x36S1
0x37S1
0x38S
0x39H

Blank cell = no glyph at the particular Unicode offset.

2. Vowels and modifiers

As Brahmic scripts are abugidas, a consonant carries the inherent a sound. This is assumed to be the default. Other phonetic properties are denoted by appending a specific modifier digit to the consonant's code (see §3). The broader keys are derived by dropping specific digit classes (see §3.1).

InputRule→ key2
ml
hi
bn
kn
inherent 'a'K
mlകി
hiकि
bnকি
knಕಿ
vowel sign (mātrā)K4
mlകു
hiकु
bnকু
knಕು
vowel signK5
mlക്ക
hiक्क
bnক্ক
knಕ್ಕ
gemination (doubled glyphs)K2
mlകം
hiकं
bnকং
knಕಂ
anusvāraK3

3. Digit codes

Phonetic details beyond the mapped letter in the table is recorded by appending a modifier digit to that letter's Roman code. The finest key, key2, retains all properties; the coarser key1 and key0 are formed by stripping specific digit classes from them.

DigitDescriptionExampleExists in
1"Hard" sound (retroflex vs. sibilant distinction)ण → N1,   ष → S1key1, key2
2Gemination (doubled consonant)ക്ക → K2key2
3Anusvāraകം → K3key0, key1, key2
4–9 Mātrā (vowel signs)
4 = i/ī
5 = u/ū
6 = e/ē
7 = ai
8 = o/ō
9 = au
കി → K4key2

0 is not in any digit classes. 0 is not a modifier digit, but is a base sound coded after the Greek θ's "th" for the dental plosives ('dantya', eg: त थ द ध). While the anusvāra 3 is a modifier, nasalisation inherently changes a word's identity too much, so it is not discarded like other digit classes. So, both these digits are retained in every key, while every other digit class ([1], 02, 4–9) are dropped in classes to form the coarser keys.

3.1. Digit class transformation

KeyTypeResult
key2Full encodingRetain all digits
key1key2 sans [2] and [4-9]Drops gemination and vowel signs. Keeps hard sounds [1] and anusvāra [3]
key0key1 sans [1]Drops hard sounds and keeps only anusvāra [3]

3.1.2. Example

The algorithm parses a word character by character, left to right, in one pass. It holds at most one consonant at a time and tracks three states (idle, holding a consonant, and inside a conjunct after a virama), deciding at each character whether to emit the held consonant, merge a group, or attach a modifier. Below is an example of the state transformation with the Malayalam word നീലക്കുയിൽ (nīlakkuyil, a cuckoo).

CharacterTypeStateCodekey2
consonantHold consonant N 
◌ീvowel sign (ī)Emit held N and then the vowelN4N4
consonantHold consonant LN4
consonantEmit held L and hold KLN4L
◌്viramaBegin conjunct while K still heldN4L
consonantSame consonant after virama, so it's a geminateK2N4LK2
◌ുvowel sign (u)Attach the vowel5N4LK25
consonantHold consonant YN4LK25
◌ിvowel sign (i)Emit held Y and then the vowelY4N4LK25Y4
chillu (ḷ)Emit the bare consonant directlyLN4LK25Y4L

The final key2 is then reduced to two coarser keys by stripping away classes of digits:

KeyHashDigits stripped
key2N4LK25Y4L
key1NLKYLVowel/gemination digits [2], [4-9]
key0NLKYLHard-sound marker [1]

Source code

See main.js for the Javascript implementation of the algorithm used in the demo of this page.