I have a task to search through several hundred files and extract duplicate detail records and keep them grouped with their header record. If no duplicate detail record exists, don't pull the header. For example, an input file could look like this:
input.txt
HA
D1
D2
D2
D3
D4
D4
HB
D1
D2
HC
D1
D1
D2
D3
D3
The output would be:
output.txt
HA
D2
D4
HC
D1
D3
Would it be possible to do this with AWK? I do not know python.
What distinguishes Header and data? Is there a fixed list of Headers or was the input file generated after pasting your several hundred files? Can you explain the exact requirements?
Thanks for the replies.
These is actually multiple files of daily extracts of expense report data from a transactional system. each file is made up of individual expense reports (header records) and the expense line items for each report (detail records). We had a situation where some detail records, but not all, were duplicated. This occurred in some output files, but not all. My requirements are to identify, by export file, the duplicate records, attached to their respective header records. We need this information to send to the system of record to correct these errors. It (hopefully) will be a one time fix. Also, I do not know perl, but am willing to learn enough to use it as a solution.
Can u explan me your script how you are printing header as i am not that much familiar in perl.Will your script print all the outputs with headers assuming there are lots of headers.
Thanks for the script. I REALLY appreciate the help!!
I do not know perl, although I am looking into thanks to the resources on this site. How would I change the code so I could run it taking a file name in as a parameter such as:
thanks for the script. When I ran it, it only returned header records, with no duplicate details. How would I supress returning headers with no duplicates.
#!/usr/bin/python
import sys,os
filename = sys.argv[1]
headers = {"HA": [] , "HB": [], "HC": []}
for line in open(filename):
line = line.strip() #strip new line
if headers.has_key(line):
flag = 1
k = line
elif flag:
headers[k].append(line)
print "File %s:" % filename
for key in sorted( headers.keys() ):
for s in set( headers[key] ):
if headers[key].count(s) > 1:
print key
print '\n'.join( set(headers[key] ) )
break