How do I execute this find command multiple times throughout many directories with the ampersand '&' command?

Viewed 287

writing a script that should utilize the find command (locate would not work because of some issue between the database and the filesystem, already tried and does not work) to locate a file by name or extension, but because the filesystem is about ~200TB, it would not be as fast and efficient to run a single find command. My idea is to run find through multiple directories with the '&' command as I believe it would be more efficient that way, although I my be wrong. my current script so far is

#!/bin/bash

echo "Enter either file name or format:"
read FileV

echo "Input the absolute path to directory"
read Dir

for d in $Dir
do
        ( cd $d && find ???
3 Answers

You can use xargs to parallize the command. Run this in the top-level directory and it will farm out find commands to as many CPUs as it can.

One advantage of doing it this way is that since you're not backgrounding the processes, you wont need to worry about the jobspec output cluttering stdout.

Change the -name part to whatever you're looking for:

for dir in */; do echo "$dir"; done | xargs -P0 -I_ find _ -type f -name "*.sh" > /tmp/outfile

From the xargs manpage

 -P max-procs, --max-procs=max-procs
              Run up to max-procs processes at a time; the default is 1.  If max-procs
              is 0, xargs will run as many processes as possible at a time.  

The bottleneck for the OP question is disk access. Given 200TB data size, only small part of the disk information will be in cached memory. As a result, the operation will be disk-bound. Running in parallel will have relatively little effect - the processes will be waiting for disk IO most of the time.

Following the proposals from other users - using locate, or similar is likely to provide more efficient search. Even a simple "Do It Youself" index - cron job that will do "find ...", and store the output in a file can be combined with grep to quickly find the files by name, and yield yield 100X speedup.

To run one find instance per subdirectory in a path, you can use:

for d in "$Dir"/*/
do
  find "$d" -name "$FileV" &
done
wait

You may also want to consider installing and enabling locate, the standard file indexing and searching feature. It will periodically index all files, and then let you search through the index much faster than re-iterating all the files again.

Related