[
https://issues.apache.org/jira/browse/HADOOP-13340?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16322663#comment-16322663
]
Ruslan Dautkhanov commented on HADOOP-13340:
--------------------------------------------
[~jlowe] A workaround might be to compress only files for which compression
makes sense? For example it doesn't make a lot of sense to compress tiny files.
It may make sense to compress when files are over a few Kb. Not sure if a
hard-coded source file size would do. If it's over a threshold, that one file
will be compressed.
> Compress Hadoop Archive output
> ------------------------------
>
> Key: HADOOP-13340
> URL: https://issues.apache.org/jira/browse/HADOOP-13340
> Project: Hadoop Common
> Issue Type: New Feature
> Components: tools
> Affects Versions: 2.5.0
> Reporter: Duc Le Tu
> Labels: features, performance
>
> Why Hadoop Archive tool cannot compress output like other map-reduce job?
> I used some options like -D mapreduce.output.fileoutputformat.compress=true
> -D
> mapreduce.output.fileoutputformat.compress.codec=org.apache.hadoop.io.compress.GzipCodec
> but it's not work. Did I wrong somewhere?
> If not, please support option for compress output of Hadoop Archive tool,
> it's very neccessary for data retention for everyone (small files problem and
> compress data).
--
This message was sent by Atlassian JIRA
(v6.4.14#64029)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]