The input is:
>HWI-EAS382_30FC7AAXX:4:1:1580:1465
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
>HWI-EAS382_30FC7AAXX:4:1:1062:1640
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
>HWI-EAS382_30FC7AAXX:4:1:272:629
AAAAAAAAGCTATAGTCTCGTCACACATACTCACAA
>HWI-EAS382_30FC7AAXX:4:1:1033:1135
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
>HWI-EAS382_30FC7AAXX:4:1:1421:27
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
My desired output is:
>HWI-EAS382_30FC7AAXX:4:1:1580:1465
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
>HWI-EAS382_30FC7AAXX:4:1:272:629
AAAAAAAAGCTATAGTCTCGTCACACATACTCACAA
What command line I should type to remove those duplicated sequence?
Thanks for all of your advise.
Hi, fajohnson...
Your command line is worked. But still left all the header of the nucleotide sequence. Do you have better idea that I just remain the first header of those same nucleotide sequence?
My input:
>HWI-EAS382_30FC7AAXX:4:1:631:449
>HWI-EAS382_30FC7AAXX:4:1:93:1407
>HWI-EAS382_30FC7AAXX:4:1:154:1123
>HWI-EAS382_30FC7AAXX:4:1:912:1008
>HWI-EAS382_30FC7AAXX:4:1:57:316
>HWI-EAS382_30FC7AAXX:4:1:1287:1193
>HWI-EAS382_30FC7AAXX:4:1:1451:1559
>HWI-EAS382_30FC7AAXX:4:1:1431:1913
TTTCCGCGAACTGCAAAAGACGTTTCGTATGCCGTT
My output just want left this:
>HWI-EAS382_30FC7AAXX:4:1:631:449
TTTCCGCGAACTGCAAAAGACGTTTCGTATGCCGTT