I�m really new at Unix and have to write a script to copy some files from some locations to another one.
I need to be sure that each file has finished copied before starting to copy the other one. I believe unix returns the control to a script before the command really ends.
I'd be sure to check the return value of each step. As a general rule, you want to check the return value of anything that does return a value. The built-in variable $? will have a 0 for success (with a few exceptions).
You mean check for every directory copied?
I used the other script because i didn't found any other way to do it.
By the way, can you please tell me how can i do to "join" several directories to a backup job?
Like this
tar cvf /device ./First #This one always cvf
tar ??? /device ./Second
...
How does it return anything and how can i treat the variable?
It would be helpful if you identify why you want to ensure a sequential order of the copies. Is this for something like an Oracle database backup where you are putting tablespaces in backup mode, then backing up member datafiles? Either way, I'd like to understand why you want to be sure that each file has finished copied before starting to copy the other one.
The cp checks can be done many ways. There is a built-in conditional that can be specified with ||, for example
cp ick1 ick9 || cp ick88 ick99 || cp ick3 success
The next operation will not be performed unless the previous is successful. The same conditional would work with tar specifying each directory. You would just have to be sure that you are specifying a non-rewinding device (ie. /dev/rmt/1mn instead of /dev/rmt/1m).
You can do this by polling the process with the ps command. If you were to put the tar command in the background then you could capture its PID. Then sit and wait for it to complete by checking to see if that PID is still running.
To detect errors, you can also do a "set -e" which will cause the shell to exit if any command returns a non zero status.
Detecting that a command failed is different from detecting that a command launched a background process before returning. It would not be easy to detect that. If a command started a separate process, it would often be to run some other command. So that "ps" would not be very general.
I have never seen a version of tar that would behave that way. The final close of the tape device may involve a rewind operation. The close system call should not return to tar until the rewind is complete. If that is not happening, the tape driver has a bug. I can't imagine any other circumstances where this might even be possible.
In general, commands return after they are finished. Any exceptions to this should be noted in the docs for the command. I really only encounter this behavior when I expect to. For instance, a command like "start-web-server" or something would reasonably be expected to leave a running web server. But the resulting process would most likely be called "httpd".
That�s the behaviour of the tar command that I should expect, but my boss insists in saying that it uses buffers and gives back control to the OS before the actual writing finishes.
I�ll try to build something using the process number and wait until it�s not running, just to please him
That is probably true. But it depends on the exact command, the exact filesystem(s), and the exact os version, and possibly on the exact hardware. It also depends on his definition of "buffer".
In the case of a tar to tape, the data will be on tape before the command exits. In case of copy disk to disk, it's complicated. Unix tries to do read-ahead and write-behind with most disk i/o to a file in a filesystem. Usually disk reads and writes go to and from a buffer cache. Some file systems now have special options to control this, (vxfio on HP-UX) but that is kinda new. A syncer program runs every so often (30 seconds on HP-UX) to flush out old data. There is a sync program that can flush the data as well. But sync and syncer only schedule the i/o... some time may be needed for the driver to process the request.
Today's disk drives are really little computers and they too have a cache. They can have a feature called immediate-report which makes them pretend that a write is finished when it has just been cached locally in the drive. A well designed drive can even withstand (in theory) a powerfail in this state. When power is restored, the drive will write the data. The OS may be able to control this behavior (scsictl in HP-UX).
One universal way to be sure that all data has been flushed from the system to the drives: unmount the filesystem. The unmount will not succeed until after the data is flushed. Well, this bummed people out when applied to nfs filesystems and some OS's now discard(!) the data on a nfs dismount if the nfs server is not responding. (Nothing is sacred anymore...)
My company is also very paranoid about losing data. We replicate our data to three massive servers... 2 in North America and one in Europe.
The script its suposed to copy files to a tape device using tar. It�s a script for one of our customers, using a SUN Ultra 10 server with Solaris 8 as OS.
As far as I knew, the disks have internal buffers, but tape devices not.
I had the chance to get some more info about the problem and the situation is this one:
The problem is (as far as I know) that the script first copies files from the current machine to the tape device and when finishes calls other server to backup some files using the same device, but the second server gets an error message stating that the device is still in use.
If you imagine something I can use to be sure that the device is free, please let me know.
(actually the script is issuing an "at" command 15 minutes ahead in order to let the tape device finish working, but they don�t like this approach).
Add a "mt -f status" to script after it finishes with drive and before it invokes the the next script. And post the results. Let the second script fail and post the error message.
Can you watch the process? Is the tape still rewinding?
Did you invoke the first script like this?
script1 > /dev/rmt/0m
If so the drive is opened by the script itself and you need to do a:
exec >&-
to close it before you invoke the second script.