[
https://issues.apache.org/jira/browse/HADOOP-8989?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13586375#comment-13586375
]
Jonathan Allen commented on HADOOP-8989:
----------------------------------------
Some quick answers to your questions ...
The CommandFactory is added into the Command so that it can be used by the
-exec expression (it needs to know what class to pass the found files into). It
seemed cleaner to do it this way than to call FsShell.run for each file. In
hindsight I think the factory should be setting it rather than FsShell (in the
same place it calls setConf).
Should be able to avoid most of the static blocks. As an aside, why are static
blocks worse than initialising static variables? (or are static variables other
than constants also bad?)
(arg == null) needs to be checked because there may be a null on the Deque,
nothing to do with args being empty or not. I'm not sure how a null would have
got on the Deque but the method shouldn't be making assumptions about that.
I needed to override processPaths so that I could implement the -depth option.
processPaths always processes the current item before recursing the directory.
It looks like somebody has added a postProcessPath method since I started this
so that might solve the problem, I'll revisit.
I should get a chance to action the comments later this week.
> hadoop dfs -find feature
> ------------------------
>
> Key: HADOOP-8989
> URL: https://issues.apache.org/jira/browse/HADOOP-8989
> Project: Hadoop Common
> Issue Type: New Feature
> Reporter: Marco Nicosia
> Assignee: Jonathan Allen
> Attachments: HADOOP-8989.patch, HADOOP-8989.patch, HADOOP-8989.patch,
> HADOOP-8989.patch, HADOOP-8989.patch, HADOOP-8989.patch, HADOOP-8989.patch,
> HADOOP-8989.patch
>
>
> Both sysadmins and users make frequent use of the unix 'find' command, but
> Hadoop has no correlate. Without this, users are writing scripts which make
> heavy use of hadoop dfs -lsr, and implementing find one-offs. I think hdfs
> -lsr is somewhat taxing on the NameNode, and a really slow experience on the
> client side. Possibly an in-NameNode find operation would be only a bit more
> taxing on the NameNode, but significantly faster from the client's point of
> view?
> The minimum set of options I can think of which would make a Hadoop find
> command generally useful is (in priority order):
> * -type (file or directory, for now)
> * -atime/-ctime-mtime (... and -creationtime?) (both + and - arguments)
> * -print0 (for piping to xargs -0)
> * -depth
> * -owner/-group (and -nouser/-nogroup)
> * -name (allowing for shell pattern, or even regex?)
> * -perm
> * -size
> One possible special case, but could possibly be really cool if it ran from
> within the NameNode:
> * -delete
> The "hadoop dfs -lsr | hadoop dfs -rm" cycle is really, really slow.
> Lower priority, some people do use operators, mostly to execute -or searches
> such as:
> * find / \(-nouser -or -nogroup\)
> Finally, I thought I'd include a link to the [Posix spec for
> find|http://www.opengroup.org/onlinepubs/009695399/utilities/find.html]
--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira