How to extract duplicate records with associated header record

All,

I have a task to search through several hundred files and extract duplicate detail records and keep them grouped with their header record. If no duplicate detail record exists, don't pull the header. For example, an input file could look like this:

input.txt
HA
D1
D2
D2
D3
D4
D4
HB
D1
D2
HC
D1
D1
D2
D3
D3

The output would be:

output.txt
HA
D2
D4
HC
D1
D3

Would it be possible to do this with AWK? I do not know python.

Thank you for your time.

What distinguishes Header and data? Is there a fixed list of Headers or was the input file generated after pasting your several hundred files? Can you explain the exact requirements?

I'd use perl for this...

$ ./input.pl 
HA
D2
D4
HC
D1
D3
$ cat ./input.pl 
#!/usr/bin/perl
# Script to print headers and duplicate items from input.txt

use warnings;
use strict;
my @records;
undef $/;

open ( INPUT, "< input.txt" ) || die "Couldn't open input file: $!\n";
# use a look-ahead assertion here
@records = split( /^(?=(?:H))/m, <INPUT> );
foreach my $record ( @records ) {
   my @lines = split( /\n/, $record );
   my $header = $lines[0];
   my %linehash;
   my $headerdone = 0;
   foreach my $line ( @lines ) {
      $linehash{$line}++;  
   }
   foreach my $key ( sort ( keys ( %linehash ) ) ) {
      my $value = $linehash{$key};
      if ( $value > 1 ) { 
         if ( $headerdone == 0 ) {
            printf( "%s\n", $header );
            $headerdone++;
         }
         printf( "%s\n", $key );
      }
   }
}
close ( INPUT );

exit ( 0 );

Cheers
ZB

Thanks for the replies.
These is actually multiple files of daily extracts of expense report data from a transactional system. each file is made up of individual expense reports (header records) and the expense line items for each report (detail records). We had a situation where some detail records, but not all, were duplicated. This occurred in some output files, but not all. My requirements are to identify, by export file, the duplicate records, attached to their respective header records. We need this information to send to the system of record to correct these errors. It (hopefully) will be a one time fix. Also, I do not know perl, but am willing to learn enough to use it as a solution.

thanks again for posting a reply.

#! /opt/third-party/bin/perl

my ($content, $i, $header, $headerprint, %fileHash);
open(FILE, "< a") || die "Unable to open file : $!\n";

while( chomp($content = <FILE>) ) {
  if( $content =~ m/^H/ ) {
    $headerprint = 0;
    $header = $content;
    %fileHash = ();
  }
  else {
    if( $headerprint == 0 ) {
      print "$header\n"; $headerprint = 1;
    }
    print "$content\n" if exists $fileHash{$content};
    $fileHash{$content} = $i++;
  }
}

exit 0

This shell script should do for you.

#! /usr/bin/ksh
r=`sort $1|uniq -d`
if [ -z "$r" ]; then
echo " No duplicate record found"
else
k=`sort -u $1`
echo "output.txt:" >>outputfile
echo "$k" >> outputfile
exit 0
fi

But this wont,

categorize the ouput based on the header as the OP requested for :slight_smile:

matrix:

Can u explan me your script how you are printing header as i am not that much familiar in perl.Will your script print all the outputs with headers assuming there are lots of headers.

Take the same input what the OP had provided .

>output

HA
D2
D4
HB
HC
D1
D3

Now try running your script and check the ouput :slight_smile:

This is my script output.

[root@localhost ~]# cat outputfile
output.txt:
D1
D2
D3
D4
HA
HB
HC

Following is the ouput that OP had requested for,
(extract from the start of the thread )

>output.txt
HA
D2
D4
HB
HC
D1
D3

under each header HA; HB; HC;
the script should display the items that are repeated! :slight_smile:

matrixmadhan,

Thanks for the script. I REALLY appreciate the help!!

I do not know perl, although I am looking into thanks to the resources on this site. How would I change the code so I could run it taking a file name in as a parameter such as:

>perlcode.pl filename.txt

#! /opt/third-party/bin/perl

my ($content, $i, $header, $headerprint, %fileHash);
open(FILE, "< ".$ARGV[0]) || die "Unable to open file : $!\n";

while( chomp($content = <FILE>) ) {
  if( $content =~ m/^H/ ) {
    $headerprint = 0;
    $header = $content;
    %fileHash = ();
  }
  else {
    if( $headerprint == 0 ) {
      print "$header\n"; $headerprint = 1;
    }
    print "$content\n" if exists $fileHash{$content};
    $fileHash{$content} = $i++;
  }
}

exit 0

anbu23,

thanks for the script. When I ran it, it only returned header records, with no duplicate details. How would I supress returning headers with no duplicates.

Slightly modified,

try this

#! /opt/third-party/bin/perl

my ($content, $i, $header, $headerprint, %fileHash);
open(FILE, "< a") || die "Unable to open file : $!\n";

while( chomp($content = <FILE>) ) {
  if( $content =~ m/^H/ ) {
    foreach $att (@disp) {
      if( $headerprint == 0 ) {
        $headerprint = 1;
        print "$header\n";
      }
      print "$att\n";
    }
    $headerprint = 0;
    $header = $content;
    %fileHash = ();
    @disp = ();
  }
  else {
      push(@disp, $content) if exists $fileHash{$content};
      $fileHash{$content} = $i++;
  }
}
foreach $att (@disp) {
  if( $headerprint == 0 ) {
    $headerprint = 1;
    print "$header\n";
  }
  print "$att\n";
}

close(FILE);
exit 0

matrixmadhan,

Thanks, it pulls the duplicates with their headers!! I REALLY appreciate all of the help I received on this issue.

Thanks to everyone who responded!!

Using awk...

$ awk '/^H/{h=$0;f=0}$0==p{print (f?"":h ORS) p;f=1}{p=$0}' file1 
HA
D2
D4
HC
D1
D3

Python alternative:

#!/usr/bin/python
import sys,os
filename = sys.argv[1]
headers = {"HA": [] , "HB": [], "HC": []}
for line in open(filename):
    line = line.strip() #strip new line
    if headers.has_key(line):
        flag = 1
        k = line
    elif flag:
        headers[k].append(line)
print "File %s:" % filename
for key in sorted( headers.keys() ):
    for s in set( headers[key] ):
        if headers[key].count(s) > 1:
            print key
            print '\n'.join( set(headers[key] )  )
            break

python script.py inputfile