[I] Remove Telugu normalization of vu వు to ma మ from IndicNormalizer [lucene]

via GitHub Tue, 13 May 2025 10:05:12 -0700


Trey314159 opened a new issue, #14659:
URL: https://github.com/apache/lucene/issues/14659


   ### Description
   
   Telugu vu వు and ma మ are visually similar—akin to English "rn" and "m"—but 
they should not be conflated. Names like వెంకటరామ (Venkatarama) and వెంకటరావు 
(Venkatarao) and words like 
[మండే](https://te.wiktionary.org/wiki/%E0%B0%AE%E0%B0%82%E0%B0%A6%E0%B0%BF) and 
[వుండే](https://te.wiktionary.org/wiki/%E0%B0%B5%E0%B1%81%E0%B0%82%E0%B0%A6%E0%B0%BF)
 (links to Telugu Wiktionary) are distinct.
   
   It's like conflating "rn" and "m" to merge _burn/bum_ and _corn/com._ It 
could happen when reading quickly or with poor handwriting, but it is not 
something that should happen for search indexing.
   
   I notice that some of the Telugu elements of IndicNormalizer are in 
TeluguNormalizer, but this mapping is not—which is good!
   
   (Sorry for the botched pull request. Obviously this change would also affect 
some tests, which need to be updated or re-evaluated.)
   
   ### Version and environment details
   
   My version:
   "distribution" : "opensearch",
   "number" : "1.3.20",
   "lucene_version : "8.10.1"
   
   Running on x86_64 GNU/Linux in Docker 4.15.0 on MacOS 13.6.3.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: issues-unsubscr...@lucene.apache.org.apache.org

For queries about this service, please contact Infrastructure at:
us...@infra.apache.org


---------------------------------------------------------------------
To unsubscribe, e-mail: issues-unsubscr...@lucene.apache.org
For additional commands, e-mail: issues-h...@lucene.apache.org

[I] Remove Telugu normalization of vu వు to ma మ from IndicNormalizer [lucene]

Reply via email to