# Sort strings with numbers

**URL:** <https://community.unix.com/t/sort-strings-with-numbers/354844>\
**Category:** Shell Programming and Scripting\
**Created:** [August 23, 2015, 2:24pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844 "2015-08-23T14:24:35Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![aydj](https://community.unix.com/user_avatar/community.unix.com/aydj/32/1795_2.png) [@aydj](https://community.unix.com/u/aydj)\
**Post date:** [August 23, 2015, 2:24pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/1 "2015-08-23T14:24:35Z")

</div>

I want to sort my data first by the 2nd field then by the first field.  
I can't use

```nohighlight
sort -V

```

because I don't have gnu sort and cannot install one.  
How do I go about this?

Input:

```nohighlight
G456 KT1 34
K234 KT10 45
L2 KT2 26
H5 LAF2 28
F3 LAF2 36

```

Output:

```nohighlight
G456 KT1 34
L2 KT2 26
K234 KT10 45
F3 LAF2 36
H5 LAF2 28

```

---

<div class="post-metadata">

**Author:** ![Don\_Cragun](https://community.unix.com/user_avatar/community.unix.com/don_cragun/32/2682_2.png) [@Don\_Cragun](https://community.unix.com/u/Don_Cragun)\
**Post date:** [August 23, 2015, 4:13pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/2 "2015-08-23T16:13:19Z")

</div>

There are some data specific ways to pre-process your data into fields that a standard `sort` utility can process, sort it, and then post-process the results to get back your original data in your desired sorted order.

For example, with your sample data (which has single spaces as field separators and the 1st two fields each starting with a string of one or more alphabetic characters followed by a string of one or more decimal digits), you could add spaces before the first digit in the 1st and 2nd fields, sort with options `-k3,3 -k4,4n -k1,1 -k2,2n` , and then remove the 3rd and 1st spaces from the sorted output.

If your data isn't as simple as shown in your sample (some data in the 1st two fields with no letters, no digits, some numbers with a leading decimal point, numbers containing more then one decimal point, more than one string of letters with numbers interspersed, etc.), then the pre-processing and post-processing steps would be correspondingly more complex.

---

<div class="post-metadata">

**Author:** ![RudiC](https://community.unix.com/letter_avatar/rudic/32/5_5575768a8748004e209b776fc1b2916d.png) [@RudiC](https://community.unix.com/u/RudiC)\
**Post date:** [August 24, 2015, 3:53am UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/3 "2015-08-24T03:53:40Z")

</div>

Without loss of generality (esp. concerning Don Cragun's comments) and taylored to your sample, this would satisfy your request:

```nohighlight
sed -r 's/(^| )([[:alpha:]]+)([[:digit:]]+)/\1\2 \3/g' file | sort -k3,3 -k4,4n -k1,1 -k2,2n | sed -r 's/([[:alpha:]]+) ([[:digit:]]+)/\1\2/g'

```

---

<div class="post-metadata">

**Author:** ![drl](https://community.unix.com/user_avatar/community.unix.com/drl/32/542_2.png) [@drl](https://community.unix.com/u/drl)\
**Post date:** [August 24, 2015, 12:13pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/4 "2015-08-24T12:13:59Z")

</div>

Hi.

Utility `msort` can handle all this internally:

```nohighlight
       msort provides twelve types of key comparison: lexicographic, numeric,
       numeric string, hybrid, by string length, by angle, by date, by domain
       name, by time, by ISO8601 date/time stamp, by month name, and random.
-- man msort

```

```nohighlight
#!/usr/bin/env bash

# @(#) s1	Demonstrate ordering of hybrid strings, msort.
# Requires: libc6 (>= 2.7-1), libgmp3c2, libicu38 (>= 3.8-5), libtre4, libuninum5
# If not in repository, see:
# homepage: http://www.billposer.org/Software/msort.html

# Utility functions: print-as-echo, print-line-with-visual-space, debug.
# export PATH="/usr/local/bin:/usr/bin:/bin"
LC_ALL=C ; LANG=C ; export LC_ALL LANG
pe() { for _i;do printf "%s" "$_i";done; printf "\n"; }
pl() { pe;pe "-----" ;pe "$*"; }
db() { ( printf " db, ";for _i;do printf "%s" "$_i";done;printf "\n" ) >&2 ; }
db() { : ; }
C=$HOME/bin/context && [-f $C] && $C

FILE=${1-data1}

pl " Input data file $FILE:"
cat $FILE

pl " Expected results:"
cat expected-results.txt

pl " Results:"
msort -q -l -n2,2 -chybrid -n1,1 -clexicographic $FILE

exit 0

```

producing:

```nohighlight
$ ./s1

Environment: LC_ALL = C, LANG = C
(Versions displayed with local utility "version")
OS, ker|rel, machine: Linux, 2.6.26-2-amd64, x86_64
Distribution : Debian 5.0.8 (lenny, workstation) 
bash GNU bash 3.2.39

-----
 Input data file data1:
G456 KT1 34
K234 KT10 45
L2 KT2 26
H5 LAF2 28
F3 LAF2 36

-----
 Expected results:
G456 KT1 34
L2 KT2 26
K234 KT10 45
F3 LAF2 36
H5 LAF2 28

-----
 Results:
G456 KT1 34
L2 KT2 26
K234 KT10 45
F3 LAF2 36
H5 LAF2 28

```

Best wishes ... cheers, drl

---

<div class="post-metadata">

**Author:** ![aydj](https://community.unix.com/user_avatar/community.unix.com/aydj/32/1795_2.png) [@aydj](https://community.unix.com/u/aydj)\
**Post date:** [August 24, 2015, 6:11pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/5 "2015-08-24T18:11:46Z")

</div>

@RudiC, I get this error,

```nohighlight
sed: illegal option -- r

```

, can awk be used?

---

<div class="post-metadata">

**Author:** ![Don\_Cragun](https://community.unix.com/user_avatar/community.unix.com/don_cragun/32/2682_2.png) [@Don\_Cragun](https://community.unix.com/u/Don_Cragun)\
**Post date:** [August 24, 2015, 6:25pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/6 "2015-08-24T18:25:18Z")

</div>

> [@aydj](#):
>
> @RudiC, I get this error,
> 
> ```plaintext
> sed: illegal option -- r
> 
> ```
> 
> , can awk be used?

What operating system are you using? It is a good idea to always give us this information (and the shell you're using) when you ask for help so we can choose utilities and options that will work in your environment when we make suggestions.

---

<div class="post-metadata">

**Author:** ![aydj](https://community.unix.com/user_avatar/community.unix.com/aydj/32/1795_2.png) [@aydj](https://community.unix.com/u/aydj)\
**Post date:** [August 24, 2015, 6:45pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/7 "2015-08-24T18:45:25Z")

</div>

OS = Solaris 10  
Shell = ksh

---

<div class="post-metadata">

**Author:** ![Scrutinizer](https://community.unix.com/user_avatar/community.unix.com/scrutinizer/32/1216_2.png) [@Scrutinizer](https://community.unix.com/u/Scrutinizer)\
**Post date:** [August 24, 2015, 10:40pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/8 "2015-08-24T22:40:24Z")

</div>

With awk you could try:

```nohighlight
nawk '{p=$0;$1=$2 FS $1; gsub(/[0-9]+/," &",$1)}{print $1,p}' file | sort -n | nawk '{print $5,$6,$7}'

```

---

<div class="post-metadata">

**Author:** ![RudiC](https://community.unix.com/letter_avatar/rudic/32/5_5575768a8748004e209b776fc1b2916d.png) [@RudiC](https://community.unix.com/u/RudiC)\
**Post date:** [August 25, 2015, 4:55am UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/9 "2015-08-25T04:55:09Z")

</div>

I don't have access to a solaris system. Does it's `sed` support the `-E` option (which is equivalent to `-r` )?

---

<div class="post-metadata">

**Author:** ![aydj](https://community.unix.com/user_avatar/community.unix.com/aydj/32/1795_2.png) [@aydj](https://community.unix.com/u/aydj)\
**Post date:** [August 25, 2015, 5:25am UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/10 "2015-08-25T05:25:53Z")

</div>

@Scrutinizer, seem not to be working, This is the output I got:

```plaintext
G456 KT1 34
K234 KT10 45
L2 KT2 26
F3 LAF2 36
H5 LAF2 28

```

---------- Post updated at 10:25 ---------- Previous update was at 10:22 ----------

> [@rudic](#):
>
> I don't have access to a solaris system. Does it's `sed` support the `-E` option (which is equivalent to `-r` )?

No -E option:

```plaintext
sed: illegal option -- E
sed: illegal option -- E

```

---

<div class="post-metadata">

**Author:** ![RudiC](https://community.unix.com/letter_avatar/rudic/32/5_5575768a8748004e209b776fc1b2916d.png) [@RudiC](https://community.unix.com/u/RudiC)\
**Post date:** [August 25, 2015, 6:47am UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/11 "2015-08-25T06:47:42Z")

</div>

try

```nohighlight
sed 's/\(^\| \)\([[:alpha:]]\+\)\([[:digit:]]\+\)/\1\2 \3/g' file | sort -k3,3 -k4,4n -k1,1 -k2,2n | sed 's/\([[:alpha:]]\+\) \([[:digit:]]\+\)/\1\2/g'

```

---

<div class="post-metadata">

**Author:** ![aydj](https://community.unix.com/user_avatar/community.unix.com/aydj/32/1795_2.png) [@aydj](https://community.unix.com/u/aydj)\
**Post date:** [August 25, 2015, 9:01am UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/12 "2015-08-25T09:01:52Z")

</div>

> [@rudic](#):
>
> try
> 
> ```plaintext
> sed 's/\(^\| \)\([[:alpha:]]\+\)\([[:digit:]]\+\)/\1\2 \3/g' file | sort -k3,3 -k4,4n -k1,1 -k2,2n | sed 's/\([[:alpha:]]\+\) \([[:digit:]]\+\)/\1\2/g'
> 
> ```

Not working, this is the output:

```plaintext
L2 KT2 26
H5 LAF2 28
G456 KT1 34
F3 LAF2 36
K234 KT10 45

```

---

<div class="post-metadata">

**Author:** ![RudiC](https://community.unix.com/letter_avatar/rudic/32/5_5575768a8748004e209b776fc1b2916d.png) [@RudiC](https://community.unix.com/u/RudiC)\
**Post date:** [August 25, 2015, 9:11am UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/13 "2015-08-25T09:11:52Z")

</div>

This is becoming a bit confusing. Try

```nohighlight
sed 's/\([[:alpha:]][[:alpha:]]*\)\([[:digit:]][[:digit:]]*\)/\1 \2/g' file | sort -k3,3 -k4,4n -k1,1 -k2,2n | sed 's/\([[:alpha:]][[:alpha:]]*\) \([[:digit:]][[:digit:]]*\)/\1\2/g'

```

---

<div class="post-metadata">

**Author:** ![Scrutinizer](https://community.unix.com/user_avatar/community.unix.com/scrutinizer/32/1216_2.png) [@Scrutinizer](https://community.unix.com/u/Scrutinizer)\
**Post date:** [August 25, 2015, 12:20pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/14 "2015-08-25T12:20:20Z")

</div>

> [@aydj](#):
>
> @Scrutinizer, seem not to be working, This is the output I got:
> 
> ```plaintext
> G456 KT1 34
> K234 KT10 45
> L2 KT2 26
> F3 LAF2 36
> H5 LAF2 28
> 
> ```
> 
> [..]

Yes, the sort wasn't right. Try:

```plaintext
nawk '{p=$0;$1=$2 FS $1; gsub(/[0-9]+/," &",$1)}{print $1,p}' file | sort -k1,1 -k2,2n -k3,3 -k4,4n | nawk '{print $5,$6,$7}'

```

---

> [@rudic](#):
>
> This is becoming a bit confusing. Try
> 
> ```plaintext
> sed 's/\([[:alpha:]][[:alpha:]]*\)\([[:digit:]][[:digit:]]*\)/\1 \2/g' file | sort -k3,3 -k4,4n -k1,1 -k2,2n | sed 's/\([[:alpha:]][[:alpha:]]*\) \([[:digit:]][[:digit:]]*\)/\1\2/g'
> 
> ```

On Solaris one needs to use `/usr/xpg4/bin/sed` to use POSIX character classes..

---

<div class="post-metadata">

**Author:** ![aydj](https://community.unix.com/user_avatar/community.unix.com/aydj/32/1795_2.png) [@aydj](https://community.unix.com/u/aydj)\
**Post date:** [August 25, 2015, 6:23pm UTC](https://community.unix.com/t/sort-strings-with-numbers/354844/15 "2015-08-25T18:23:07Z")

</div>

> [@scrutinizer](#):
>
> Yes, the sort wasn't right. Try:
> 
> ```plaintext
> nawk '{p=$0;$1=$2 FS $1; gsub(/[0-9]+/," &",$1)}{print $1,p}' file | sort -k1,1 -k2,2n -k3,3 -k4,4n | nawk '{print $5,$6,$7}'
> 
> ```
> 
> ---
> 
> On Solaris one needs to use `/usr/xpg4/bin/sed` to use POSIX character classes..

Thanks, it works!
