Tip: Find all shell scripts in a folder

I was looking for a way to find all shell script in my system (or in a folder) and I found nothing that works good. Just one old topic from 2014 with very limited use expecting .sh or the file to be executable (bash scripts can be sourced and do not need to always have executable permissions), looking for the text in the first line, etc.

Finally I found a solution for me that is working just fine by using the file analysis command:

find ~+ -type f -exec file {} \; | grep "shell script" | awk -F : '{ print $1 }'

Maybe it can be useful also for someone else, even if you change the grep text to find some other type of files.

Nice! I put "tip" in the title.
A bit shortened, and generalized for standard shells:

find "$PWD" -type f -exec file {} + | awk -F: '/shell script/{print $1}'

This is really nifty ... and useful. Apparently it requires proper shebangs in all shell scripts, right?

Yes, it looks to be reported as a bash script the cmd file requires a shebang to be present. The name of the script is not important, no need to be *.sh or similar. No need to be also executable.
I needed to collect this way all bash/sh/csh... scripts and to cat all of them in a single file, then to search the file for examples how to use some commands that I learn at the moment. There is a lot to learn from the professionals but first I had to find where all these script files are hidden. 90+ MB total of sources found from my Ubuntu installation.

also, worth a [re]read of documentation on the file command man file

Man pages definitely are the main source of information but I'm talking about real life examples rarely shown there. Like this one for example:

printf "\x6f\x6b" | dd of=$orgfile bs=1 seek=58724780 count=2 conv=notrunc &>> /dev/null

Immediately we can see what it is doing but not so obvious it can be used this way for the beginners looking only at man printf/dd. At least for me, learning from tested examples is much easier way to learn.

Hi, I wasn't implying otherwise. Just that knowing how the file command works is also germane.
as for the example , not sure how that's relevant wrt determining if a file is a [shell/python/perl/awk/.../] script or not.
!# on line one is what I'd be using, everything else is open to interpretation

maybe 'script' and 'executable' to cover other interpreters .

a contrived example (on macos)

file *
python.bash:  a /bin/python script text executable, ASCII text
script:       Bourne-Again shell script text executable, ASCII text
t.js:         a /usr/local/bin/node script text executable, ASCII text
t.randomname: Perl script text executable
test.sh:      Bourne-Again shell script text executable, ASCII text
verifyScript:        ASCII text # a script, but no #! .... 
x.j:          a /bin/rubbish script text executable, ASCII text

Disclaimer: I'm the current author of the rawhide (rh(1)) program referred to in this response (at github/raforg).

TL;DR Just scroll to the end

If you are only looking for actual shell scripts, which are executable, this rawhide (rh(1)) command will do it faster than find(1) plus a separate file(1) process for each candidate (because it doesn't require any file(1) processes, or grep(1), or awk(1)):

rh 'f && ix && ("text/plain*".mime || "*sh*script*".what)'

It lists regular files that are executable, whose mimetype starts with text/plain, or whose file(1) output would contain "sh script" or "shell script".

It handles very plain shell scripts with no #!/bin/sh line (on the understanding that any text file with no #! line that is executable is treated by the operating system as a /bin/sh script when it is executed), and it handles shell scripts with #!/bin/sh or #!/bin/ksh or #!/bin/zsh etc., and it handles shell scripts with #!/usr/bin/env zsh etc.

It can be made faster by limiting the expensive analysis (I/O) just to files below a certain size by adding "&& sz<50K" immediately after the "&& ix" (for the find(1) version, add "-size -50k" before "-exec"). That will eliminate I/O on files that are too large to be shell scripts. How large is too large? You decide. I've got one that's 63KiB!

If you need to include non-executable files that are sourced into a shell script (i.e., shell fragments), I don't know what to do. file(1) (and rh(1)) see them as text/plain and ASCII text (at least on macOS and Linux). I could recommend making all shell fragments executable, but that would be cheating. :slight_smile:

The find(1) command at the top of this discussion doesn't report non-executable shell fragments either (for me). In fact, it doesn't report a lot of things. My test cases had:

1 no #! line and not executable [therefore not a shell script]
2 no #! line but executable
3 #!/bin/sh
4 #!/bin/ksh
5 #!/bin/zsh
6 #!/bin/bash
7 #!/usr/bin/env ksh
8 #!/usr/bin/env zsh
9 #!/usr/bin/env bash

This rh(1) command listed all but #1. The find(1) command at the top of this discussion only reported #3 #4 #6 and #9. It missed #2 because it had no #! line. It missed #5 because the file(1) output contained "zsh script" rather than "shell script". It missed #7 and #8 because the file(1) output contained "/usr/bin/env ksh script" and "/usr/bin/env zsh script" (i.e., file(1) just quoted the #! line), but it did report #9 because file(1) detects "/usr/bin/env bash" as a special case.

To include scripts in other languages like perl, python, js, etc., it's slightly simpler:

rh 'f && ix && ("text/plain*".mime || "*script*".what)'

Unfortunately, the following isn't as good (it doesn't report scripts without a #! line):

rh 'f && ix && "*script*".what'

But that might be enough for your needs, if all of your scripts have a #! line.

Here are the find(1) equivalents (different on Linux and macOS) of the less good version (that misses scripts without a #! line).

GNU/Linux version:

find . -type f -perm /111 -exec file '{}' ';' | awk -F: '/script/ { print $1 }'

macOS version:

find . -type f -perm +111 -exec file '{}' ';' | awk -F: '/script/ { print $1 }'

Here are the find(1) equivalents of the complete version (that does detect shell scripts without a #! line):

GNU/Linux version:

find . -type f -perm /111 -exec /bin/sh -c '(file --mime-type "{}"; file "{}") | grep -Eq "text/plain|script"' ';' -print

macOS version:

find . -type f -perm +111 -exec /bin/sh -c '(file --mime-type "{}"; file "{}") | grep -Eq "text/plain|script"' ';' -print

The find(1) command at the start of this discussion does detect non-executable shell fragments if they start with #!/bin/sh, but that #! line is non-functional. Without the execute bit, the kernel will never use it to execute the file. It can only be a shell fragment, and the #! line is probably there for syntax-highlighting purposes in a text editor, rather than for execution. But if your non-executable shell fragments do have a #! line for syntax-highlighting purposes, they can be included with this command:

rh 'f && ("*script*".what || ix && "text/plain*".mime)'

If you're happy to exclude actual executable shell scripts without a #! line, this will do:

rh 'f && "*script*".what'

To limit it to just shell scripts (i.e., no perl, python, etc.):

rh 'f && "*sh*script*".what'

These super short versions aren't the most accurate, but they're probably fine.

Thanks for pointing out the shortcomings of the GNU/Linux file command.
(BTW it is even worse in the commercial Unixes. At least Solaris makes a half-assed attempt to support it.)
Here is a little improvement:

find "$PWD" -type f -exec file {} + | awk -F: '/(shell|[a-z]sh) script/{print $1}'

Sure the rh tool is more specialized in such tasks.
If you frequently examine your files then go for it. The link is GitHub - raforg/rawhide: find files using pretty C expressions

A slightly shorter but equivalent version is:

find "$PWD" -type f -exec file {} + | awk -F: '/sh(ell)? script/{print $1}'

But that excludes actual executable shell scripts that don't have a #! line, but which are interpreted by the kernel as /bin/sh scripts if and when they are executed.

To include those, you would need (GNU/Linux version - for macOS replace /111 with +111):

find "$PWD" -type f \( -exec /bin/sh -c 'file "{}" | grep -q sh.*script' ';' -o -perm /111 -exec /bin/sh -c 'file --mime-type "{}" | grep -q text/plain' ';' \) -print

I definitely prefer the rh(1) equivalent :slight_smile: :

rh 'f && ("*sh*script*".what || ix && "text/plain*".mime)'

P.S. I don't think rh(1) is specialized. It's just a file-finding program like GNU find(1) (which I agree is definitely much better than all the commercial versions of find(1)).

Disclaimer: I'm a beginner in Bash and just wanted to share my joy of my own successful one-liner finding, and I clearly understand that it is not perfect and it is based on how good is the tool file itself.

Before that I did a script in the traditional logical way knowing that bash has parameter -n to not execute the script, just to verify the syntax. False positive was empty file which was easy to fix, but also .pem files are successfully validated as bash script (no idea why):

#!/bin/bash
shopt -s dotglob
for f in $PWD/*
do
	if [ -f $f ]; then
		if [ -s $f ]; then
			bash -n $f 2> /dev/null
			if [ $? -eq 0 ]; then
				echo $f
			fi
		fi
	fi
done

Please don't help me to make it on one line, I'll figure it out soon, just for exercise, no practical use of that version.

So, thanks for deeper analyzing the topic. My initial intention was to collect some scripts for learning, not absolutely all of them at any price.

PS: I installed rh (V3.2 compiled from sources) on my Ubuntu 22.04 for testing and started with the provided examples here. Unfortunately nothing is working:

test@U2204LTS:~$ rh 'f && x && ("text/plain*".mime || "*sh* script*".what)'
rh: command line: -e 'f && x && ("text/plain*".mime || "*sh* script*".what)': line 1 byte 6: undefined identifier: identifier x
test@U2204LTS:~$ rh 'f && ("text/plain*".mime || "*sh* script*".what)'
rh: command line: -e 'f && ("text/plain*".mime || "*sh* script*".what)': line 1 byte 24: invalid string suffix: .mime (expected pattern modifier or reference file field)

The nested ifs are like logical ANDs.
With &&

	if [ -f "$f" ] && [ -s "$f" ] && file "$f" | grep -wq script && bash -n "$f" 2> /dev/null
	then
		echo "$f"
	fi

The remaining if can be replaced as well: you can chain a && echo "$f"

Oops. Sorry about that. The missing x identifier comes from the additional brevity config mentioned in the language documentation (man rawhide.conf). It's defined as:

x { executable }

You could put that in your ~/.rhrc but it isn't essential. The above commands can be changed to use ix rather than x which is the same, but it's in the default config (/etc/rawhide.conf). I'll change the above commands to use ix instead of x.

The .mime and .what pattern modifiers are missing because your system didn't have the libmagic library and header files when rh was compiled. For Debian/Ubuntu, install the libmagic-dev package, then reconfigure, rebuild and reinstall rawhide. Also make sure that you have pcre2 installed (libpcre2-dev package on Debian/Ubuntu).

libmagic-dev was missing and I added the suggested aliases in ~/.rhrc, libpcre2-dev was already installed. Now everything is working. Thanks!
Scanned and found over 33,000 script files in my system by

rh 'f && ("*sh*script*".what || ix && "text/plain*".mime)'

Very few of them are with missing shebang that are found as executable text. Most of them are coming created by Windows through Samba share folders providing wrong executable permissions for all files. For me it is better to use:

rh 'f && "*sh*script*".what'

This command does seem to be slightly obscure in any case.

(a) "\x6f\x6b" is just 'ok', so no need for the hex version.

(b) seek=58724780 is mysterious, especially as orgfile is a variable (and unquoted !).

Such a magic number deserves a variable name, and to show how it might be determined.

I am guessing that this is merely appending the new text, so the option oflag=append would then be appropriate, instead of (perhaps) some prior code to find the current size with stat (or even ls or (horror) wc -c).

(c) Setting bs=1 makes dd rather slow -- it seems to inhibit block buffering.

The option oflag=seek_bytes changes the behaviour of seek=nnn, such that it uses the default block buffering, but still seeks in bytes (not blocks).

The default bs is 512 bytes: I would suggest setting bs=4096 or whatever your file system block size happens to be (for large transfers, bs=1M).

You can concatenate like oflag=append,seek_bytes. I suspect multiple oflag operands are also permitted, but the man page does not explicitly say so.

(d) I think count=2 is not required, as the input length is also determined by the pipe reaching EOF.

If you changed the input to 'perfect', you might miss the need to also change count=7.

(e) I would probably suppress statistics with status=none, and not redirect actual error messages to /dev/null. No data will come out on stdout if you have of=xxx, so you probably want to see stdout (if any) too.

The moral: essentially, it can often be unwise to trust existing scripts too far.

That was just one line from a larger script generated by a custom diff tool in order to patch one executable after copying it from a read only to a temp location before execution. Yes, the selected shortest line by me is indeed just the text "ok" but the idea was to patch any type of bytes and on the rest of the lines there were mostly binary data. I changed the input file name to something else, it was hard-coded real filename generated by the script. You quote a parameter variable when you don't know what filename it will hold. In this case the filename is known, on read only location that will not change and it was just "/main.elf", no spaces that will make the difference and can stay unquoted. BS=1 according to my understanding is because they patch bytes and require offset in the file in bytes. If it was the default 512 or something else, the skip will skip that match block size elements and will not be able to start from a specific file byte. Deleting the dd output from so many lines (dd makes output for statistics at the end and for errors if missing permissions, etc) was according to the logic of the script to judge the final patch result with md5 checksum at the end of the patched file, not from the result of each script line. It may be not a perfect example, but I liked the way to use generated binary data from the script to be used with dd, not using a real file for the input.

I fully quote all expansions like "${myVar}". The rationale for this is:

(a) I generally leave my production code on client sites, where it gets maintained by people with less knowledge than myself (that's why they pay me to do it). I get fewer annoying (and unproductive) call-backs if my code is consistently fully quoted, rather than rely on users to understand the implicit restrictions. That is why shellcheck flags all instances.

(b) Most of the bash substitution options only work when in braces, so it is easier to brace everything than get it wrong sometimes. Neither $k//o/x nor "$k//o/x" do what "${k//o/x}" does.

Your understanding of the interaction between dd seek and bs is incomplete.

In GNU/dd, the options bs=4096 seek=12977 oflag=seek_bytes and bs=1 seek=12977 have identical results. The only difference (checked by strace) is that the first version makes 2 system calls (reading 5 bytes "fi zz" from fd0 and writing 5 bytes to fd1), and the second version makes 10 system calls of 1 byte each. Obviously not significant where you are flicking a few bytes, but the usual use of dd is to deal with bulk data, where the overhead would be serious (by orders of magnitude) without optimal buffering.

I don't see relying on the final md5 as very helpful. Sure, it tells you one (or more) of the patches failed, but not which ones, so you end up re-running them manually to find the problem.

Thanks for the deeper dive on this example. There is always something to learn. strace appears to be too complex for me as a beginner but shellcheck is something new for me that looks to be extremely valuable tool. Thanks!