My problem is that I wish to identify the suffixes ( i.e. the possible identical longest strings moving from right to left) which are adjoined to the names and store such strings along with their frequency in a separate file with the following conditions
the suffix string should be at least between 3 and 5 characters in length
the suffix string should be repeated at least 10 times in the database.
Thus in the sample given above, the script would identify only the following suffixes along with their frequency
singh 12
bai 12
The suffix
kumar
Will not be identified since it is less than 10 times.
I had posted the query earlier, but at present I have tried to refine it with conditional constraints so that hopefully only the most pertinent suffixes will be identified. There could be a few false positives but I could weed them out.
I work in a Windows environment and PERL or AWK script would be helpful.
Many thanks and all good wishes for the New Year to all the folks who take their valuable time off to help people solve their problems
Many thanks. It worked very well. When I posted the request, I knew that there are chances of false positives, but a list of suffixes is easier to handle than wading through thousands of lines.
I can also tweak the awk script if I wish to set the range
Happy New Year and thanks once more
While the earlier method worked and I had to tweak a few suffixes manually, I have been rethinking the process of identification of suffixed names and after going through nearly 40 to 50 thousand names, I have identified a pattern. Very often, in nearly 95% of the cases,the name that is suffixed is also a name by itself as in the example below and comes first in my rev sort followed by names to which it is suffixed.
Could it be possible to extract such suffixes given that the suffix is a stand-alone name as in the case of
kumar
bai
singh
with the proviso that the standalone name is suffixed at least three times to another name. This would obviate the need for blind search and also false positives. I know that this could possibly miss out a few suffixes, but from my analysis, this could provide a more accurate solution.
Would it be possible to devise a PERL or AWK script to identify such cases.
Many thanks once again for all kind help.
Thanks a lot. It works well, all I had to do was trim off short words from the list and which in no way were suffixes, and I managed to get a pretty comprehensive lst of suffixes.
I have been studying the syntax of the script and there is one part which perplexes me. The rest I could grab
NR == 1
Could you please explain what this really does.
Thanks once again and a Happy New Year.
To explain why Scrutinizer made that suggestion, note that the expression $0 !~ IX "$" looks for a line that does not have the contents of the variable IX immediately followed by the end of the input line.
Using the expression $0 !~ IX "^" looks for a line that does not have the contents of the variable IX followed by the start of a line (which will ALWAYS happen). Try:
awk '
{if ($0 !~ "^" IX || NR == 1) IX = $0
else CNT[IX]++
}
END {for (c in CNT) print c, CNT[c]
}
' file
which (with your sample input in the file named file ) produces the output:
Thanks to both Scrutinizer and Don Cragun not only for the code but for taking pains to explain the placement of the operator and why it behaves in that manner.
Thanks alot