Regular Expression

Hello,

I want to extract text between <td class="di_resultscolumnheader"> and </td>.
I wrote the below code to extract text. But I am able to extract the text for the first match only.
Can some one help me in this?

Thanks in advance.

Code:

if ($line =~ /<td class="di_resultscolumnheader">(.*?)<\/td>/g)
{
	print $1,"\n";
}

Sample Input:

<td class="di_resultscolumnheader">Object</td><td class="di_resultscolumnheader">Precipitant</td><td class="di_resultscolumnheader">Therapeutic�Class</td>

Sample Output:

Object
Precipitant
Therapeutic�Class

You could try something like this:

sed 's!<td class="di_resultscolumnheader">.*</td>!&!g; s!</td>!\n!g; s!<td class="di_resultscolumnheader">!!g;s/\n$//; /^$/d; ' input-file >output-file

This doesn't handle the case where there might be text between </td> and <td...> . A simple, but slightly longer, sed would be needed to deal with that.

$ nawk -F"[<>]" '{for(i=1;i<=NF;i++)if($i~/resultscolumnheader/){print $(i+1)}}' input.txt
Object
Precipitant
Therapeutic�Class
# sed -n '/^<td class="di_resultscolumnheader">/,/<\/td>/{s///;s/<td class="[^ ]*">//g;s//\n/g;:a;s/<\/td>/\n/;/td.*td/ta;s/<\/td>$//;p}' infile
Object
Precipitant
Therapeutic�Class
awk '/^td class="di_resultscolumnheader"/{print $2}' RS=\< FS=\> infile

Perl:

while (/<td class="di_resultscolumnheader">(.*?)<\/td>/g)
{
    print "$1\n";
}