Use records from one file to delete records in another file

file_in_1:

1 2 3 4
5 6 7 8
9 10 11 12
13 14 15 16
17 18 19 20
21 22 23 24
25 26 27 28
29 30 31 32

file_in_2:

9 10 11 12
21 22 23 24
1 2 3 4
17 18 19 20

file_out:

5 6 7 8
13 14 15 16
25 26 27 28
29 30 31 32

How can I use each record in file file_in_2 and delete the
corresponding records in file file_in_1 and produce
the file file_out.

Note:
My files can have thousands, even millions of records.
and tend to be not sorted.

Thank you,
Kenny.

fgrep -vxf file_in_2 file_in_1 >file_out

--- if your fgrep has the -f option. Otherwise, perhaps something along the lines of

sed -e 's/[][\\.*$^]/\\&/g' -e 's!.*!/^&$/d!' file_in_2 |
sed -f - file_in_1 >file_out

If your sed doesn't grok -f - either, you need to put the output from the first sed script in a temporary file.

sed -e 's/[][\\.*$^]/\\&/g' -e 's!.*!/^&$/d!' file_in_2 >temporary
sed -f temporary file_in_1 >file_out
rm temporary

Different versions of sed understand slightly different variants of regular expression syntax, so there may be a need to adjust the first sed script slightly. But if your input file is just numbers and spaces, that is absolutely nothing to worry about.

As a last resort, maybe Perl could be useful:

perl -nle 'if ($. == ++$l) { $r{$_} = 1; close ARGV if eof; next}
print unless $r{$_}' file_in_2 file_in_1 >file.out

If you really have millions of lines in both sets, it might be fruitful to import them into a database or something if your regular line-oriented tools choke on really large files.

Thanks era,

The fgrep doesn't produce the desired result.
I think it might require the input files to be sorted first.

Your first sed code works.
In my example files you can see that file 2 is unsorted.

Thanks again,
Kenny.

fgrep doesn't care about sort order. The examples you posted work here with fgrep. Perhaps there is some minor whitespace difference or control character somewhere ...?

awk >file_out 'NR==FNR{_[$0];next}!($0 in _)' file_in_2 file_in_1

Use nawk or /usr/xpg4/bin/awk on Solaris.

what about the following?
Yes, it could take a while; yes, it could all be done in one command (done this way to show what is happening); yes, it does re-organize the data.

> cat ifile1 ifile2 | sort -n >ifile3
> cat ifile3
1 2 3 4
1 2 3 4
5 6 7 8
9 10 11 12
9 10 11 12
13 14 15 16
17 18 19 20
17 18 19 20
21 22 23 24
21 22 23 24
25 26 27 28
29 30 31 32

then the command for uniq

> cat ifile3 | uniq -u 
5 6 7 8
13 14 15 16
25 26 27 28
29 30 31 32