# How to find a phrase and pull all lines that follow until the phrase occurs again?

**URL:** <https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639>\
**Category:** Shell Programming and Scripting\
**Created:** [August 28, 2013, 2:48pm UTC](https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639 "2013-08-28T14:48:15Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Scottie1954](https://community.unix.com/letter_avatar/scottie1954/32/5_5575768a8748004e209b776fc1b2916d.png) [@Scottie1954](https://community.unix.com/u/Scottie1954)\
**Post date:** [August 28, 2013, 2:48pm UTC](https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639/1 "2013-08-28T14:48:15Z")

</div>

I want to burst a report by using the page number value in the report header. Each section starts with

```nohighlight
*PAGE NO:* 1

```

Each section might have several pages, but the next section always starts back at 1.

So I want to find the "_PAGE NO:_ 1" value and pull all lines that follow until "_PAGE NO:_ 1" appears again.

I've tried using awk search for the phrase and pull newline characters but I can't rely on a set number of lines that follow the phrase.

Thank you!

-Scottie1954

---

<div class="post-metadata">

**Author:** ![in2nix4life](https://community.unix.com/user_avatar/community.unix.com/in2nix4life/32/709_2.png) [@in2nix4life](https://community.unix.com/u/in2nix4life)\
**Post date:** [August 28, 2013, 3:03pm UTC](https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639/2 "2013-08-28T15:03:30Z")

</div>

I don't know the format of the file you're trying to parse or any other info, but this sed command can grab the lines between two given patterns (excluding said patterns):

```nohighlight
sed -n '/<pattern1>/,/<pattern2>/{//!p};'

```

---

<div class="post-metadata">

**Author:** ![Corona688](https://community.unix.com/letter_avatar/corona688/32/5_5575768a8748004e209b776fc1b2916d.png) [@Corona688](https://community.unix.com/u/Corona688)\
**Post date:** [August 28, 2013, 3:54pm UTC](https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639/3 "2013-08-28T15:54:42Z")

</div>

```nohighlight
awk '/PAGE NO:\* 1$/ { P=!P ; next } P' inputfile > outputfile

```

---

<div class="post-metadata">

**Author:** ![Scottie1954](https://community.unix.com/letter_avatar/scottie1954/32/5_5575768a8748004e209b776fc1b2916d.png) [@Scottie1954](https://community.unix.com/u/Scottie1954)\
**Post date:** [August 28, 2013, 4:45pm UTC](https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639/4 "2013-08-28T16:45:49Z")

</div>

Yeah, I've been a bit cryptic with what's in the file as it has mostly confidential data. It would take some effort to redact. Anyway, I've tried both replies here with no luck. The sed suggestion produces a parsing error and the awk returns no data. I'm running /usr/bin/sh on hp-ux.

The number that I want to search between is in the last column of the report header. If I use the following code I can see the numbers change, but I don't know how to grab the lines in between. Thank you.

```nohighlight
grep 'PAGE NO:' reportfile | awk '{ print $NF }'

```

---

<div class="post-metadata">

**Author:** ![Corona688](https://community.unix.com/letter_avatar/corona688/32/5_5575768a8748004e209b776fc1b2916d.png) [@Corona688](https://community.unix.com/u/Corona688)\
**Post date:** [August 28, 2013, 6:20pm UTC](https://community.unix.com/t/how-to-find-a-phrase-and-pull-all-lines-that-follow-until-the-phrase-occurs-again/333639/5 "2013-08-28T18:20:49Z")

</div>

You probably just need to tweak the awk version's regex until it starts matching the "page number" lines. We can't do that for you, since you can't post the data...
