I want to extract text between <td class="di_resultscolumnheader"> and </td>.
I wrote the below code to extract text. But I am able to extract the text for the first match only.
Can some one help me in this?
Thanks in advance.
Code:
if ($line =~ /<td class="di_resultscolumnheader">(.*?)<\/td>/g)
{
print $1,"\n";
}
sed 's!<td class="di_resultscolumnheader">.*</td>!&!g; s!</td>!\n!g; s!<td class="di_resultscolumnheader">!!g;s/\n$//; /^$/d; ' input-file >output-file
This doesn't handle the case where there might be text between </td> and <td...> . A simple, but slightly longer, sed would be needed to deal with that.