Help with editing string elements

Hi All
I have a question.
I would like to edit some string characters by replacing with characters of choice located in another file. For example in sample file

>S5_SK1.chr01
NNNNNNNNNNNNNNNNNNNCAGCATGCAATAAGGTGACATAGATATACCCACACACCACACCCTAACACTAACCCTAATCTAACCCTGGCCAACCTGTTT
CTCAACTTACCCTCCATTACCCTACCTCCACTCGTTACCCTGTCCCATTCAACCATACCACTCCGAACCACCATCCATCCCTCTACTTACTACCACTCAC

excluding the first line beginning with ">" I would like to change string elements at position 10 which is N to T, at position 100 which is T to C and at position 157 which is an A to G in the above example based on information present in the following file:

chr01    10    T   
chr01    100   C    
chr01    157   G    

So my expected output would be

>S5_SK1.chr01
NNNNNNNNNTNNNNNNNNNCAGCATGCAATAAGGTGACATAGATATACCCACACACCACACCCTAACACTAACCCTAATCTAACCCTGGCCAACCTGTTC
CTCAACTTACCCTCCATTACCCTACCTCCACTCGTTACCCTGTCCCATTCAACCATGCCACTCCGAACCACCATCCATCCCTCTACTTACTACCACTCAC

Can anyone suggest how I can do this, preferably using Perl as I'm learning that. I would like to perform this over multiple files and thus I'll really appreciate your input.
Hv a nice day:)
Cheers

Try:

#!/usr/bin/perl
open I,"position_file";
@b=<I>;
%h=map {((split /\s+/)[1],(split /\s+/)[2]);} @b;
local $/;
open J,"main_file";
$_=<J>;
s/(>.*)//;
$n=$1;
s/\n//g;
@F=split //,$_;
for $i (keys %h){
  $F[$i-1]=$h{$i}
}
$_=join "", @F;
s/.{100}/$&\n/g;
print "$n\n";
print "$_\n";

Let me know if something needs explanation.

while read a a b
do
     sed "s/./$b/$a" infile >infile.tmp
     cat infile.tmp >infile
done <file2
rm infile.tmp

Where file2 is the file containing the

chr01 10 T  
chr01 100 C
...

stuff.

@Bartus:
Hi Bartus I'm not getting any output with this apart from two blank lines !! :frowning:
Any reason why that might be happening. My example main file is

>S5_SK1.chr01
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNCAGCATGCAATAAGGTGACATAGATATACCCACACACCACACCCTAACACTAACCCTAATCTAACCCTGGCCAACCTGTTT
CTCAACTTACCCTCCATTACCCTACCTCCACTCGTTACCCTGTCCCATTCAACCATACCACTCCGAACCACCATCCATCCCTCTACTTACTACCACTCAC
CCACCGTTACCCTCCAATTACCCATATCCAACTCCACTGCCACTTACCCTGCCATTCCTCTACCATCCACCATCTGCTACTCACTGTACTGTTGTTCTAC

and position file is

chr01   62      C
chr01   63      A
chr01   70      G
chr01   80      C
chr01   100     A
chr01   200     T
chr01   300     G
chr01   400     C
chr01   550     A
chr01   599     A

Cheers and thanks for looking into my question
Hv a nice day :slight_smile:

Weird... It is working for me:

oracle@solaris:~/unix/dna$ cat file
>S5_SK1.chr01
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNCAGCATGCAATAAGGTGACATAGATATACCCACACACCACACCCTAACACTAACCCTAATCTAACCCTGGCCAACCTGTTT
CTCAACTTACCCTCCATTACCCTACCTCCACTCGTTACCCTGTCCCATTCAACCATACCACTCCGAACCACCATCCATCCCTCTACTTACTACCACTCAC
CCACCGTTACCCTCCAATTACCCATATCCAACTCCACTGCCACTTACCCTGCCATTCCTCTACCATCCACCATCTGCTACTCACTGTACTGTTGTTCTAC
oracle@solaris:~/unix/dna$ cat pos
chr01   62      C
chr01   63      A
chr01   70      G
chr01   80      C
chr01   100     A
chr01   200     T
chr01   300     G
chr01   400     C
chr01   550     A
chr01   599     A
oracle@solaris:~/unix/dna$ ./a.pl
>S5_SK1.chr01
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNCANNNNNNGNNNNNNNNNCNNNNNNNNNNNNNNNNNNNA
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNT
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNG
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNC
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNN
NNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNANNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNNAN
NNNNNNNNNNNNNNNNNNNCAGCATGCAATAAGGTGACATAGATATACCCACACACCACACCCTAACACTAACCCTAATCTAACCCTGGCCAACCTGTTT
CTCAACTTACCCTCCATTACCCTACCTCCACTCGTTACCCTGTCCCATTCAACCATACCACTCCGAACCACCATCCATCCCTCTACTTACTACCACTCAC
CCACCGTTACCCTCCAATTACCCATATCCAACTCCACTGCCACTTACCCTGCCATTCCTCTACCATCCACCATCTGCTACTCACTGTACTGTTGTTCTAC
oracle@solaris:~/unix/dna$ cat a.pl
#!/usr/bin/perl
open I,"pos";
@b=<I>;
%h=map {((split /\s+/)[1],(split /\s+/)[2]);} @b;
local $/;
open J,"file";
$_=<J>;
s/(>.*)//;
$n=$1;
s/\n//g;
@F=split //,$_;
for $i (keys %h){
  $F[$i-1]=$h{$i}
}
$_=join "", @F;
s/.{100}/$&\n/g;
print "$n\n";
print "$_\n";

Hi Bartus,
I understood why it was not workinig ... my file names were different from that in the script and so I changed it to

$ARGV[0] for position_file and $ARGV[1]

for main_file in the script and ran

./script.pl position_file main_file

and this worked :slight_smile:
Cheers n hv a nice day ahead !!
Thanks

---------- Post updated at 11:46 AM ---------- Previous update was at 11:26 AM ----------

@Bartus:
Could you also explain the map part of the code the the use of Perl in built variables throughout the code? I'll appreciate that.
Cheers :slight_smile:

map function is used to create hash from the contents of @b array (which contains lines from postion_file). When map is executed, its code is run for each element of the @b array. When processing those elements, they are assigned into "$_" variable inside "map" block, so (split /\s+/)[1] and (split /\s+/)[2] operate on them, extracting second and third field (position and letter) from each line. Those two elements are then returned by map into %h hash, populating it with position as the key and the letter begin its value.

About second part of your question, which built-in variables do you have in mind?

Thanks Bartus ... your explanation was clear ... actually the $/ variable you have already described in another post so its fine ... Could you remind me again the $& variable please? I guess I can find it on perlvar too, but in this context will be nice to hear from you.
Related to this question what if instead of replacing the characters I had to insert or delete the position_file letters in the original file ?? What could be changed in the code to do that ?? I'm imagining a push/pop function somewhere but could you enlighten me on this when you have a few moments?
Cheers and thanks for all your help always.
Hv a nice day:)

$& holds string that was matched by last regex, so in s/.{100}/$&\n/g; it will contain hundred characters, to which newline is then appended. It is done to restore your formating of hundred chars per line. As for the changes in behavior of the script, to delete characters on positions stored in "position file", substitute

$F[$i-1]=$h{$i}

with

$F[$i-1]=""

To insert characters at specified positions, substitute the same line with:

$F[$i-2].=$h{$i}

Thanks a lot Bartus
Hv a nice day and a good weekend ahead :slight_smile:
Cheers !!