Help with file processing

I have a fixed width file coming from source system. The total characters on each record is 786.

I am getting records in the file with less than and greater than this number and the process is failing bcs of this length of the record

Is there a way to spit out the bad records (length <>786) into a bad file and process the good records?

I am on linux.

Is the fixed width in bytes or in characters?
What character set are you using?
What shell are you using?
What operating system are you using?

On systems with a version of awk that conforms to the standards:

awk 'length($0) != 786' file

will print all lines in file that do not contain 786 characters (not counting the terminating <newline> character). If you want the count to include the <newline> character, change the 786 in that script to 785 .

On some systems, the awk length() function incorrectly counts bytes instead of counting characters. If the character set you're using contains multi-byte characters (such as UTF-8), the number of bytes in a line may vary even though the number of characters is constant.

Please see my responses below

Is the fixed width in bytes or in characters?

What character set are you using?

What shell are you using?

What operating system are you using?

#Allowing the new line to be included in the 786 characters total.
perl -ne 'length != 786 and print' file_to_process > bad_lines_file

#Removing new lines from the 786 characters total.
perl -nle 'length != 786 and print' file_to_process > bad_lines_file

Incredibly, both BSD awk and mawk suffer from this :eek: , which is clearly a bug (mawk is supposed to be SUS v2 conformant)..

gawk and /usr/xpg4/bin/awk on Solaris functions correctly though (nawk does not, but it is not POSIX compliant so I would not expect it to).

Mac OS X picks up most of its utilities from BSD, and most of the OS X utilities do adhere to the POSIX standards requirements even in places where BSD utilities sometimes fail. The awk utility, however, fails on at least two counts:

  1. length() should count characters; but instead it counts bytes, and
  2. awk -v variable=value 'script' and awk -vvariable=value 'script' should be treated exactly the same, but the 1st form works and the 2nd form gives a syntax error.

At least both bash and ksh on OS X correctly count characters (not bytes) with ${#variable} .