awk data subsets manipulation

Hi,
I'm working on a data file with the following structure

val1,val2,flag
214.7332983,979.0259,1
12.87435571,205.7679,1
1.365976384,19.01616,1
44.08584096,205.7679,2
7.034721792,383.8778,2
189.5685503,979.0259,2
1.96352032,19.01616,2
[...]

where the field 'flag' identifies different groups. I'd like to obtain statistics on each group and save it in an output file. I'm not an expert user and it is not clear to me if it is possible to tell awk to take the first 3 lines, calculate the relevant stat (piping the first two columns of the first three lines to another shell command), print the output and move to the following group (last four lines).
Any help? Thanks in advance.

Can you post the expected result?

What I have in mind is to do the following

zcat temp.csv.gz | gawk -F ','  '{if($3 == 1) print $3,$5}' | STAT_CMD

where STAT_CMD produces a statistics on the first two columns and the value '1' is dynamically replaced by the third field in the temp file, grouping lines according to the value of the flag.
In my example the output will be two numbers reporting the output STAT_CMD (ex. the correlation between the two) applied on these two pairs of columns (identified by the flag)

214.7332983,979.0259
12.87435571,205.7679
1.365976384,19.01616

and

44.08584096,205.7679
7.034721792,383.8778
189.5685503,979.0259
1.96352032,19.01616

Sorry if I'm not super clear.

You need to get the list of groups.

awk -F, 'NR>1 {print $3}' temp.csv | sort -u | while read group
do
  awk -v g=$group -F, '$3==g {print $1,$2}' temp.csv | STAT_CMD
done

regarding redirection to file, It depends on how the command "STAT_CMD" process your data.

I might not fully understand you.

Thanks a lot, it works nicely since my STAT_CMD reads the stdin. Only a minor modification

awk -F, 'NR>1 {print $3}' temp.csv | uniq | while read group
do
  awk -v g=$group -F, '$3==g {print $1,$2}' temp.csv | STAT_CMD
done

However it's very slow, and I have millions of lines. Any suggestions?

It will indeed be a bit slow like this. What is in STAT_CMD ?

STAT_CMD it's C program to calculate the Spearman rank correlation coefficient between two columns. I've checked and it seems that even without piping the output of gawk to my STAT_CMD it remains slow.

Thanks again for your help

So, you mean to say "temp.csv" is already sorted on field 3?
uniq works ONLY on sorted input.

Yep, it's already sorted.

Try:

awk -F, 'NR==1{next} $3!=p{close(cmd); p=$3} {print $1,$2 | cmd}' cmd=STAT_CMD temp.csv

Wow elegant and super fast,
Thanks a lot