I would like to convert the most frequent and second most frequent duplet in each row to 1 and -1 respectively ...and everything else to 0. please assist
A duplet is only AA , CC, GG and TT
- C1 C2 C3 C4 C5
R1 AA AA - - CC
R2 AC AA AA CC CC
R3 AT AT TT TT TT
R5 AT TT AA AA AA
awk 'NR>1{ for (i=2;i<=NF;i++) { if ( substr($i,1,1)==substr($i,2) ) {x++ ; for (c in x ) { if ( c > max ) c=max ; else if ( c < max ) max2=c } else $i="0"} END {print max, max2}' file
What happens if you have TWO strings with the SAME frequency? (as in line R2)
What happens if you have TWO strings which are both highest and the next to highest frequencies? ( like on line R3)
A bit verbose - probably can be done without sorting, but.... awk -f rita.awk myFile where rita.awk is:
function quicksort(data, left, right, i, last)
{
if (left >= right) # do nothing if array contains fewer
return # than two elements
quicksort_swap(data, left, int((left + right) / 2))
last = left
for (i = left + 1; i <= right; i++)
if (count[data]<count[data])
quicksort_swap(data, ++last, i)
quicksort_swap(data, left, last)
quicksort(data, left, last - 1, less_than)
quicksort(data, last + 1, right, less_than)
}
# quicksort_swap --- helper function for quicksort, should really be inline
function quicksort_swap(data, i, j, temp)
{
temp = data
data = data[j]
data[j] = temp
}
BEGIN {
split("AA,TT,GG,CC", tA,",")
for(i=1;i in tA;i++)
goodA[tA]
}
FNR==1 {print;next}
{
split("",arr)
split("",count)
tally=0
for(i=2;i<=NF;i++) {
if (!($i in goodA)) continue
if (!($i in count)) arr[++tally]=$i
count[$i]++
}
quicksort(arr,1,tally)
printf $1
for(i=2;i<=NF;i++) {
if ($i == arr[tally])
$i=1
else if ($i == arr[tally-1])
$i=-1
else $i=0
printf("%s%d%s", OFS, $i, (i==NF)?ORS:"")
}
}
This produces the following based on your sample input:
@Scrutinizer: Nice, but that has got a problem with larger duplicate count later in the line. For a line like
R1 AA AA AA AA - - CC CC TT TT TT
it yields
R1 1 1 1 1 0 0 -1 -1 0 0 0
Small modification
awk '
BEGIN {A["AA"]; A["CC"]; A["GG"]; A["TT"] }
NR>1 {mx2key=""; max=max2=0
for(i=2; i<=NF; i++) if ($i in A) { A[$i]++ }
for (i in A) {if(max<A) {max2=max
mx2key=maxkey
max=A
maxkey=i
}
else if (max2<A) {max2=A; mx2key=i}
A=""
}
for (i=2; i<=NF; i++) $i=($i==mx2key)?-1:($i==maxkey)?1:0
}
1
' file
Hi RudiC, but I think the former output is correct, no? CC has the lowest frequency on the line (less than TT) , so it should get -1 ...
--
OK I see, the OP said the second most frequent duplet, I must have misread..