Conditional identification of suffixes moving from right to left: revisited

Dear all,
I have a large database of names which I have sorted on reverse with a Perl Script. A sample is provided below

agarsingh
aghansingh
akalsingh
akamsingh
akbareesingh
akhamisingh
akramysingh
akuvsingh
anchalusingh
andaroosingh
angadsingh
anjawsingh
angibai
angobai
angurbai
angureebai
anjabai
anjiqbai
anjsbai
anjanybai
anjatibai
anjopbai
anjusbai
ankikbai
akhileshkumar
akleskumar
akshaykumar
anchalkumar
anjanikumar
ankitkumar
antimkumar

My problem is that I wish to identify the suffixes ( i.e. the possible identical longest strings moving from right to left) which are adjoined to the names and store such strings along with their frequency in a separate file with the following conditions
the suffix string should be at least between 3 and 5 characters in length
the suffix string should be repeated at least 10 times in the database.

Thus in the sample given above, the script would identify only the following suffixes along with their frequency

singh	12
bai	12

The suffix

kumar

Will not be identified since it is less than 10 times.
I had posted the query earlier, but at present I have tried to refine it with conditional constraints so that hopefully only the most pertinent suffixes will be identified. There could be a few false positives but I could weed them out.
I work in a Windows environment and PERL or AWK script would be helpful.
Many thanks and all good wishes for the New Year to all the folks who take their valuable time off to help people solve their problems

For starters, this:

$ awk '{while($1=substr($1,2)) if (length($1)>=3) A[$1]++} END{for(i in A) if(A>=10) print i, A}' file 
ngh 12
ingh 12
bai 12
singh 12

would produce the names with some false positive substrings .

Under Windows you probably need to put the script in a script file

{
  while($1=substr($1,2)) if (length($1)>=3) A[$1]++
} 
END {
  for(i in A) if(A>=10) print i, A
} 

and run it as

awk -f script_file inputfile

Many thanks. It worked very well. When I posted the request, I knew that there are chances of false positives, but a list of suffixes is easier to handle than wading through thousands of lines.
I can also tweak the awk script if I wish to set the range
Happy New Year and thanks once more

While the earlier method worked and I had to tweak a few suffixes manually, I have been rethinking the process of identification of suffixed names and after going through nearly 40 to 50 thousand names, I have identified a pattern. Very often, in nearly 95% of the cases,the name that is suffixed is also a name by itself as in the example below and comes first in my rev sort followed by names to which it is suffixed.

singh
agarsingh
aghansingh
akalsingh
akamsingh
akbareesingh
akhamisingh
akramysingh
akuvsingh
anchalusingh
andaroosingh
angadsingh
anjawsingh
bai
angibai
angobai
angurbai
angureebai
anjabai
anjiqbai
anjsbai
anjanybai
anjatibai
anjopbai
anjusbai
ankikbai
kumar
akhileshkumar
akleskumar
akshaykumar
anchalkumar
anjanikumar
ankitkumar
antimkumar

Could it be possible to extract such suffixes given that the suffix is a stand-alone name as in the case of

kumar
bai
singh

with the proviso that the standalone name is suffixed at least three times to another name. This would obviate the need for blind search and also false positives. I know that this could possibly miss out a few suffixes, but from my analysis, this could provide a more accurate solution.
Would it be possible to devise a PERL or AWK script to identify such cases.
Many thanks once again for all kind help.

How about

awk '
        {if ($0 !~ IX "$" || NR == 1) IX = $0
         else CNT[IX]++
        }
END     {for (c in CNT) print c, CNT[c]
        }
' file
kumar 7
bai 12
singh 12

Thanks a lot. It works well, all I had to do was trim off short words from the list and which in no way were suffixes, and I managed to get a pretty comprehensive lst of suffixes.
I have been studying the syntax of the script and there is one part which perplexes me. The rest I could grab

NR == 1

Could you please explain what this really does.
Thanks once again and a Happy New Year.

NR is the record counter, so this condition is true on the first line of the input stream.

Thanks a lot. Got it. I wasn't too sure hence I asked.
Once again, thank you. At my age (66) learning a new language/script takes some time.

RudiC kindly helped me to identify suffixes in a file with the following code:

{if ($0 !~ IX "$" || NR == 1) IX = $0
         else CNT[IX]++
        }
END     {for (c in CNT) print c, CNT[c]
        }

I tried to manipulate the code so that it would look for words at beginning of the string which match. Example

ram
ramvilas
ramkumar
ramsingh

and would identify

ram

as a prefixal element which is prefixed to a large number of words, the condition being that is stands by itself in the name list.
I replaced the

$

operator with a

^

to show the beginning of line, but could not identify

ram

as a prefixal element. It just spewed out an empty file.
Did I read the script wrong ? Any help with comments would be useful. Thanks

Hi, try:

$0 !~ "^" IX

instead of

$0 !~ IX "$"

To explain why Scrutinizer made that suggestion, note that the expression $0 !~ IX "$" looks for a line that does not have the contents of the variable IX immediately followed by the end of the input line.

Using the expression $0 !~ IX "^" looks for a line that does not have the contents of the variable IX followed by the start of a line (which will ALWAYS happen). Try:

awk '
{if ($0 !~ "^" IX || NR == 1) IX = $0
         else CNT[IX]++
        }
END     {for (c in CNT) print c, CNT[c]
        }
' file

which (with your sample input in the file named file ) produces the output:

ram 3

Thanks to both Scrutinizer and Don Cragun not only for the code but for taking pains to explain the placement of the operator and why it behaves in that manner.
Thanks alot