[jira] [Commented] (LUCENE-9663) Adding compression to terms dict from SortedSet/Sorted DocValues

Jaison.Bi (Jira) Thu, 14 Jan 2021 01:42:10 -0800


    [ 
https://issues.apache.org/jira/browse/LUCENE-9663?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17264719#comment-17264719
 ]


Jaison.Bi commented on LUCENE-9663:
-----------------------------------

Thanks, Adrien Grand.
{quote}My intuition is that it would actually be better to do LZ4 in addition 
to prefix compression, like we do for the terms dictionary of the inverted index
{quote}
I have compared the results between prefix + lz4 and lz4 only, and also tried 
to change the doc size per doc to see the difference. See the below result: 
||compression type||docs per block||*.dvd file size||write time cost||merge 
time cost||
|prefix + lz4|256|1.04GB|648456ms|375966ms|
|lz4-only|256|1.08GB|639489ms|350477ms|
|lz4-only|64|1.15GB|625797ms|298093ms|
|lz4-only|128|1.1GB|618034ms|320740ms|
|lz4-only|512|1.07GB|639892ms|458737ms|

It seems prefix compression + lz4 does not make significant improvement.  I 
think because the "common prefix" could be well-handled by lz4 :-) 

> Adding compression to terms dict from SortedSet/Sorted DocValues
> ----------------------------------------------------------------
>
>                 Key: LUCENE-9663
>                 URL: https://issues.apache.org/jira/browse/LUCENE-9663
>             Project: Lucene - Core
>          Issue Type: Improvement
>          Components: core/codecs
>            Reporter: Jaison.Bi
>            Priority: Trivial
>
> Elasticsearch keyword field uses SortedSet DocValues. In our applications, 
> “keyword” is the most frequently used field type.
>  LUCENE-7081 has done prefix-compression for docvalues terms dict. We can do 
> better by replacing prefix-compression with LZ4. In one of our application, 
> the dvd files were ~41% smaller with this change(from 1.95 GB to 1.15 GB).
>  I've done simple tests based on the real application data, comparing the 
> write/merge time cost, and the on-disk *.dvd file size(after merge into 1 
> segment).
> || ||Before||After||
> |Write time cost(ms)|591972|618200|
> |Merge time cost(ms)|270661|294663|
> |*.dvd file size(GB)|1.95|1.15|
> This feature is only for the high-cardinality fields. 
>  I'm doing the benchmark test based on luceneutil. Will attach the report and 
> patch after the test.



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

---------------------------------------------------------------------
To unsubscribe, e-mail: issues-unsubscr...@lucene.apache.org
For additional commands, e-mail: issues-h...@lucene.apache.org

[jira] [Commented] (LUCENE-9663) Adding compression to terms dict from SortedSet/Sorted DocValues

Reply via email to