Help with file manipulation

Dear All,
I have a question. So for the following sample file I would like to collect information about entries in $F[0], $F[1] & $F[4] so as to acheive the following output as shown below

SK1.chr01    854    levure5    A    G    225    .    DP=407;AF1=0.5;CI95=0.5,0.5;DP4=142,103,72,68;MQ=31;FQ=225;PV4=0.24,1,1,1    GT:PL:GQ;telomere;ID=TEL01L;Name=TEL01L
SK1.chr01    854    levure6    A    G    199    .    DP=360;AF1=0.5;CI95=0.5,0.5;DP4=138,80,70,48;MQ=31;FQ=202;PV4=0.48,1,1,1    GT:PL:GQ;telomere;ID=TEL01L;Name=TEL01L
SK1.chr01    854    levure7    A    G    225    .    DP=163;AF1=0.5;CI95=0.5,0.5;DP4=60,30,37,19;MQ=31;FQ=225;PV4=1,0.26,1,1    GT:PL:GQ;telomere;ID=TEL01L;Name=TEL01L
SK1.chr01    854    levure8    A    G    225    .    DP=194;AF1=0.5;CI95=0.5,0.5;DP4=66,46,40,28;MQ=31;FQ=225;PV4=1,1,1,1    GT:PL:GQ;telomere;ID=TEL01L;Name=TEL01L
SK1.chr01    12745    levure5    C    G    185    .    DP=125;AF1=0.5;CI95=0.5,0.5;DP4=45,21,25,20;MQ=23;FQ=155;PV4=0.23,1,1,1    GT:PL:GQ
SK1.chr01    12745    levure6    C    G    197    .    DP=85;AF1=0.5;CI95=0.5,0.5;DP4=30,15,18,18;MQ=23;FQ=153;PV4=0.17,1,1,1    GT:PL:GQ
SK1.chr01    12745    levure7    C    G    152    .    DP=36;AF1=0.5;CI95=0.5,0.5;DP4=10,7,11,6;MQ=22;FQ=42;PV4=1,1,1,1    GT:PL:GQ
SK1.chr01    12745    levure8    C    G    173    .    DP=63;AF1=0.5;CI95=0.5,0.5;DP4=21,16,12,14;MQ=23;FQ=98;PV4=0.45,1,1,1    GT:PL:GQ
SK1.chr02    16511    levure5    G    A    148    .    DP=43;AF1=1;CI95=1,1;DP4=2,1,16,19;MQ=24;FQ=-85;PV4=0.59,5.9e-05,1,1    GT:PL:GQ
SK1.chr02    16511    levure6    G    A    127    .    DP=35;AF1=0.5;CI95=0.5,0.5;DP4=4,3,7,16;MQ=25;FQ=30;PV4=0.37,0.0035,0.24,1    GT:PL:GQ

Expected output:

chr01    854     AAAA     GGGG
chr01    12745   CCCC     GGGG
chr02    16511   GG       AA

Could someone help me figure out a way to do this?
Cheers and hv a nice day:)

nawk '{x=$1;sub(".*[.]","",x);A[$2]=x;B[$2]=B[$2]$4""}END{for(i in A) print A,i,B}' inputfile

Ooops, i missed the $5 here you go :

nawk '{x=$1;sub(".*[.]","",x);A[$2]=x;B[$2]=B[$2]$4"";C[$2]=C[$2]$5""}END{for(i in A) print A,i,B,C}' inputfile

Thanks ctsgnb
nawk was not supported but awk version worked :slight_smile:
Cheers

---------- Post updated at 10:28 AM ---------- Previous update was at 10:25 AM ----------

Could you comment on the code please and if you have a Perl version I'll appreciate that too :slight_smile:

Perl solution:

perl -alne '$F[0]=~s/.*\.//;push @{$a{$F[1]}},$F[3];push @{$b{$F[1]}},$F[4];$c{$F[1]}=$F[0];END{for $i (keys %a){print "$c{$i}\t$i\t",@{$a{$i}},"\t",@{$b{$i}}}}' file

Thanks Bartus :slight_smile:
Hv a nice day

Thanks for the comments ctsgnb
Hv a nice day:)

Note that the output when scanning associative array will not be ordered, so you can then | sort <whatever_options> if needed ...