Pick up successive lines containing keywords in order The 2019 Stack Overflow Developer Survey Results Are In Unicorn Meta Zoo #1: Why another podcast? Announcing the arrival of Valued Associate #679: Cesar Manara 2019 Community Moderator Election ResultsSingle record of a file getting splitted over multiple linesPick columns from a variable length csv fileBash to join columns from multiple filesFind files that contain multiple keywords anywhere in the fileText file containing filenames and hashes - extracting lines with duplicate hashesHow to cat all lines together in file/for all files in a directoryLooking for way to move even lines to the beginning of odd linescopy lines where a character occurs even number of timeschange and manipulate lines in a file using awkCompare two text files, extract matching rows of file2 plus additional rows

should truth entail possible truth

Button changing its text & action. Good or terrible?

Deal with toxic manager when you can't quit

Can the Right Ascension and Argument of Perigee of a spacecraft's orbit keep varying by themselves with time?

Do ℕ, mathbbN, BbbN, symbbN effectively differ, and is there a "canonical" specification of the naturals?

How to read αἱμύλιος or when to aspirate

Do warforged have souls?

Example of compact Riemannian manifold with only one geodesic.

What is the padding with red substance inside of steak packaging?

How to handle characters who are more educated than the author?

Did the UK government pay "millions and millions of dollars" to try to snag Julian Assange?

1960s short story making fun of James Bond-style spy fiction

"is" operation returns false with ndarray.data attribute, even though two array objects have same id

How did the audience guess the pentatonic scale in Bobby McFerrin's presentation?

The following signatures were invalid: EXPKEYSIG 1397BC53640DB551

Why are PDP-7-style microprogrammed instructions out of vogue?

Can we generate random numbers using irrational numbers like π and e?

Is it ok to offer lower paid work as a trial period before negotiating for a full-time job?

What information about me do stores get via my credit card?

How to determine omitted units in a publication

What can I do if neighbor is blocking my solar panels intentionally?

Was credit for the black hole image misappropriated?

What was the last x86 CPU that did not have the x87 floating-point unit built in?

Are there continuous functions who are the same in an interval but differ in at least one other point?



Pick up successive lines containing keywords in order



The 2019 Stack Overflow Developer Survey Results Are In
Unicorn Meta Zoo #1: Why another podcast?
Announcing the arrival of Valued Associate #679: Cesar Manara
2019 Community Moderator Election ResultsSingle record of a file getting splitted over multiple linesPick columns from a variable length csv fileBash to join columns from multiple filesFind files that contain multiple keywords anywhere in the fileText file containing filenames and hashes - extracting lines with duplicate hashesHow to cat all lines together in file/for all files in a directoryLooking for way to move even lines to the beginning of odd linescopy lines where a character occurs even number of timeschange and manipulate lines in a file using awkCompare two text files, extract matching rows of file2 plus additional rows



.everyoneloves__top-leaderboard:empty,.everyoneloves__mid-leaderboard:empty,.everyoneloves__bot-mid-leaderboard:empty margin-bottom:0;








0















I have a tab-separated file that looks as follows:



$ cat file
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I need to pick up successive lines that contain the keywords 'polyketide synthase', 'methyltransferase', and 'oxidoreductase' in that order, and write each of these sets into separate files for further analysis.



In this case, the input file would yield 2 output files which would look as follows:



$ cat file_1
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]

$ cat file_2
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I am having a hard time doing this using awk. Any suggestions?



P.S. I have other input files that contain variable number of instances of the keywords in successive lines. This is where I am getting stuck.










share|improve this question
























  • What make those two output files file_1 & file_2 different fro each other? what other files you are talking about other files that contain variable number of instances of the keywords in successive lines? please edit your question and make it a little more clear.

    – αғsнιη
    yesterday











  • @αғsнιη Sorry if I was unclear in my question. file_1 and file_2 would contain different sets of the keyword instances in successive lines (you could look at the intended output files in the question for further clarification). Also, I have made the requested edit in the question.

    – BhushanDhamale
    21 hours ago


















0















I have a tab-separated file that looks as follows:



$ cat file
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I need to pick up successive lines that contain the keywords 'polyketide synthase', 'methyltransferase', and 'oxidoreductase' in that order, and write each of these sets into separate files for further analysis.



In this case, the input file would yield 2 output files which would look as follows:



$ cat file_1
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]

$ cat file_2
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I am having a hard time doing this using awk. Any suggestions?



P.S. I have other input files that contain variable number of instances of the keywords in successive lines. This is where I am getting stuck.










share|improve this question
























  • What make those two output files file_1 & file_2 different fro each other? what other files you are talking about other files that contain variable number of instances of the keywords in successive lines? please edit your question and make it a little more clear.

    – αғsнιη
    yesterday











  • @αғsнιη Sorry if I was unclear in my question. file_1 and file_2 would contain different sets of the keyword instances in successive lines (you could look at the intended output files in the question for further clarification). Also, I have made the requested edit in the question.

    – BhushanDhamale
    21 hours ago














0












0








0








I have a tab-separated file that looks as follows:



$ cat file
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I need to pick up successive lines that contain the keywords 'polyketide synthase', 'methyltransferase', and 'oxidoreductase' in that order, and write each of these sets into separate files for further analysis.



In this case, the input file would yield 2 output files which would look as follows:



$ cat file_1
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]

$ cat file_2
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I am having a hard time doing this using awk. Any suggestions?



P.S. I have other input files that contain variable number of instances of the keywords in successive lines. This is where I am getting stuck.










share|improve this question
















I have a tab-separated file that looks as follows:



$ cat file
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I need to pick up successive lines that contain the keywords 'polyketide synthase', 'methyltransferase', and 'oxidoreductase' in that order, and write each of these sets into separate files for further analysis.



In this case, the input file would yield 2 output files which would look as follows:



$ cat file_1
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558474.1 1159543 1160595 -4330977 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558475.1 1160607 1161116 12 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011558476.1 1161113 1162129 -3 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]

$ cat file_2
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559726.1 2496640 2497560 1334511 polyketide synthase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011559727.1 2497568 2498122 8 isoprenylcysteine carboxyl methyltransferase [Mycobacterium]
GCF_000015405.1_ASM1540v1.dist_nbr_anntn WP_011562574.1 5526997 5528142 3028875 NAD(P)/FAD-dependent oxidoreductase [Mycobacterium]


I am having a hard time doing this using awk. Any suggestions?



P.S. I have other input files that contain variable number of instances of the keywords in successive lines. This is where I am getting stuck.







text-processing awk






share|improve this question















share|improve this question













share|improve this question




share|improve this question








edited 21 hours ago







BhushanDhamale

















asked yesterday









BhushanDhamaleBhushanDhamale

1664




1664












  • What make those two output files file_1 & file_2 different fro each other? what other files you are talking about other files that contain variable number of instances of the keywords in successive lines? please edit your question and make it a little more clear.

    – αғsнιη
    yesterday











  • @αғsнιη Sorry if I was unclear in my question. file_1 and file_2 would contain different sets of the keyword instances in successive lines (you could look at the intended output files in the question for further clarification). Also, I have made the requested edit in the question.

    – BhushanDhamale
    21 hours ago


















  • What make those two output files file_1 & file_2 different fro each other? what other files you are talking about other files that contain variable number of instances of the keywords in successive lines? please edit your question and make it a little more clear.

    – αғsнιη
    yesterday











  • @αғsнιη Sorry if I was unclear in my question. file_1 and file_2 would contain different sets of the keyword instances in successive lines (you could look at the intended output files in the question for further clarification). Also, I have made the requested edit in the question.

    – BhushanDhamale
    21 hours ago

















What make those two output files file_1 & file_2 different fro each other? what other files you are talking about other files that contain variable number of instances of the keywords in successive lines? please edit your question and make it a little more clear.

– αғsнιη
yesterday





What make those two output files file_1 & file_2 different fro each other? what other files you are talking about other files that contain variable number of instances of the keywords in successive lines? please edit your question and make it a little more clear.

– αғsнιη
yesterday













@αғsнιη Sorry if I was unclear in my question. file_1 and file_2 would contain different sets of the keyword instances in successive lines (you could look at the intended output files in the question for further clarification). Also, I have made the requested edit in the question.

– BhushanDhamale
21 hours ago






@αғsнιη Sorry if I was unclear in my question. file_1 and file_2 would contain different sets of the keyword instances in successive lines (you could look at the intended output files in the question for further clarification). Also, I have made the requested edit in the question.

– BhushanDhamale
21 hours ago











2 Answers
2






active

oldest

votes


















1














You can change what you are searching for as the script progresses and change where you write to each time you cycle through your terms



awk 'BEGIN 
result_file = 1;
term_id = 1;
search_terms[1] = "polyketide synthase";
search_terms[2] = "methyltransferase";
search_terms[3] = "oxidoreductase"

$0 ~ search_terms[term_id]
print $0 >> FILENAME "_" result_file;
term_id = term_id + 1;
if (term_id > 3)
result_file = result_file + 1;
term_id = 1

' input_file


This will write to input_file_1, input_file_2...






share|improve this answer
































    1














    You might test the following code, where I split your keywords into an awk array named keys with N elements. everything starts with keys[1] and we set up a flag to check the next 1 to N-1 lines if they matches the corresponding values in the array keys [index from 2 to N], any mismatches before the N-1 line will reset this flag, if it reaches the N-1 line, then all are good for output (we also reset flag=0 here so a consecutive run of flag==1 never exceeds N-1 lines):



    $ cat t24.awk
    BEGIN
    FS = OFS = "t";
    keywords = "polyketide synthase,methyltransferase,oxidoreductase";
    N = split(keywords, keys, ",")


    # flag==1 means we are doing regex_match the next N-1 lines
    # against corresponding array element in keys from [2:N]
    # once a unmatched found, turn off flag immediately
    # if the flag==1 reached N-1 lines, then print the good match
    flag
    if($NF ~ keys[NR - start_line + 1])
    F = F ORS $0;
    if (NR == start_line+N-1) print F > "out_" f++; flag = 0
    next
    else
    flag = 0;



    # set up the flag/start_line and reset F
    $NF ~ keys[1] flag = 1; F = $0; start_line= NR;


    Run the above code with awk -f t24.awk file.txt. You can set up keywords (comma delimited) from your shell(instead of hard-coded in the BEGIN block), and then use -v keywords="..." to make it more flexible.






    share|improve this answer























      Your Answer








      StackExchange.ready(function()
      var channelOptions =
      tags: "".split(" "),
      id: "106"
      ;
      initTagRenderer("".split(" "), "".split(" "), channelOptions);

      StackExchange.using("externalEditor", function()
      // Have to fire editor after snippets, if snippets enabled
      if (StackExchange.settings.snippets.snippetsEnabled)
      StackExchange.using("snippets", function()
      createEditor();
      );

      else
      createEditor();

      );

      function createEditor()
      StackExchange.prepareEditor(
      heartbeatType: 'answer',
      autoActivateHeartbeat: false,
      convertImagesToLinks: false,
      noModals: true,
      showLowRepImageUploadWarning: true,
      reputationToPostImages: null,
      bindNavPrevention: true,
      postfix: "",
      imageUploader:
      brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
      contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
      allowUrls: true
      ,
      onDemand: true,
      discardSelector: ".discard-answer"
      ,immediatelyShowMarkdownHelp:true
      );



      );













      draft saved

      draft discarded


















      StackExchange.ready(
      function ()
      StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f511906%2fpick-up-successive-lines-containing-keywords-in-order%23new-answer', 'question_page');

      );

      Post as a guest















      Required, but never shown

























      2 Answers
      2






      active

      oldest

      votes








      2 Answers
      2






      active

      oldest

      votes









      active

      oldest

      votes






      active

      oldest

      votes









      1














      You can change what you are searching for as the script progresses and change where you write to each time you cycle through your terms



      awk 'BEGIN 
      result_file = 1;
      term_id = 1;
      search_terms[1] = "polyketide synthase";
      search_terms[2] = "methyltransferase";
      search_terms[3] = "oxidoreductase"

      $0 ~ search_terms[term_id]
      print $0 >> FILENAME "_" result_file;
      term_id = term_id + 1;
      if (term_id > 3)
      result_file = result_file + 1;
      term_id = 1

      ' input_file


      This will write to input_file_1, input_file_2...






      share|improve this answer





























        1














        You can change what you are searching for as the script progresses and change where you write to each time you cycle through your terms



        awk 'BEGIN 
        result_file = 1;
        term_id = 1;
        search_terms[1] = "polyketide synthase";
        search_terms[2] = "methyltransferase";
        search_terms[3] = "oxidoreductase"

        $0 ~ search_terms[term_id]
        print $0 >> FILENAME "_" result_file;
        term_id = term_id + 1;
        if (term_id > 3)
        result_file = result_file + 1;
        term_id = 1

        ' input_file


        This will write to input_file_1, input_file_2...






        share|improve this answer



























          1












          1








          1







          You can change what you are searching for as the script progresses and change where you write to each time you cycle through your terms



          awk 'BEGIN 
          result_file = 1;
          term_id = 1;
          search_terms[1] = "polyketide synthase";
          search_terms[2] = "methyltransferase";
          search_terms[3] = "oxidoreductase"

          $0 ~ search_terms[term_id]
          print $0 >> FILENAME "_" result_file;
          term_id = term_id + 1;
          if (term_id > 3)
          result_file = result_file + 1;
          term_id = 1

          ' input_file


          This will write to input_file_1, input_file_2...






          share|improve this answer















          You can change what you are searching for as the script progresses and change where you write to each time you cycle through your terms



          awk 'BEGIN 
          result_file = 1;
          term_id = 1;
          search_terms[1] = "polyketide synthase";
          search_terms[2] = "methyltransferase";
          search_terms[3] = "oxidoreductase"

          $0 ~ search_terms[term_id]
          print $0 >> FILENAME "_" result_file;
          term_id = term_id + 1;
          if (term_id > 3)
          result_file = result_file + 1;
          term_id = 1

          ' input_file


          This will write to input_file_1, input_file_2...







          share|improve this answer














          share|improve this answer



          share|improve this answer








          edited 19 hours ago

























          answered yesterday









          Philip CoulingPhilip Couling

          2,5791123




          2,5791123























              1














              You might test the following code, where I split your keywords into an awk array named keys with N elements. everything starts with keys[1] and we set up a flag to check the next 1 to N-1 lines if they matches the corresponding values in the array keys [index from 2 to N], any mismatches before the N-1 line will reset this flag, if it reaches the N-1 line, then all are good for output (we also reset flag=0 here so a consecutive run of flag==1 never exceeds N-1 lines):



              $ cat t24.awk
              BEGIN
              FS = OFS = "t";
              keywords = "polyketide synthase,methyltransferase,oxidoreductase";
              N = split(keywords, keys, ",")


              # flag==1 means we are doing regex_match the next N-1 lines
              # against corresponding array element in keys from [2:N]
              # once a unmatched found, turn off flag immediately
              # if the flag==1 reached N-1 lines, then print the good match
              flag
              if($NF ~ keys[NR - start_line + 1])
              F = F ORS $0;
              if (NR == start_line+N-1) print F > "out_" f++; flag = 0
              next
              else
              flag = 0;



              # set up the flag/start_line and reset F
              $NF ~ keys[1] flag = 1; F = $0; start_line= NR;


              Run the above code with awk -f t24.awk file.txt. You can set up keywords (comma delimited) from your shell(instead of hard-coded in the BEGIN block), and then use -v keywords="..." to make it more flexible.






              share|improve this answer



























                1














                You might test the following code, where I split your keywords into an awk array named keys with N elements. everything starts with keys[1] and we set up a flag to check the next 1 to N-1 lines if they matches the corresponding values in the array keys [index from 2 to N], any mismatches before the N-1 line will reset this flag, if it reaches the N-1 line, then all are good for output (we also reset flag=0 here so a consecutive run of flag==1 never exceeds N-1 lines):



                $ cat t24.awk
                BEGIN
                FS = OFS = "t";
                keywords = "polyketide synthase,methyltransferase,oxidoreductase";
                N = split(keywords, keys, ",")


                # flag==1 means we are doing regex_match the next N-1 lines
                # against corresponding array element in keys from [2:N]
                # once a unmatched found, turn off flag immediately
                # if the flag==1 reached N-1 lines, then print the good match
                flag
                if($NF ~ keys[NR - start_line + 1])
                F = F ORS $0;
                if (NR == start_line+N-1) print F > "out_" f++; flag = 0
                next
                else
                flag = 0;



                # set up the flag/start_line and reset F
                $NF ~ keys[1] flag = 1; F = $0; start_line= NR;


                Run the above code with awk -f t24.awk file.txt. You can set up keywords (comma delimited) from your shell(instead of hard-coded in the BEGIN block), and then use -v keywords="..." to make it more flexible.






                share|improve this answer

























                  1












                  1








                  1







                  You might test the following code, where I split your keywords into an awk array named keys with N elements. everything starts with keys[1] and we set up a flag to check the next 1 to N-1 lines if they matches the corresponding values in the array keys [index from 2 to N], any mismatches before the N-1 line will reset this flag, if it reaches the N-1 line, then all are good for output (we also reset flag=0 here so a consecutive run of flag==1 never exceeds N-1 lines):



                  $ cat t24.awk
                  BEGIN
                  FS = OFS = "t";
                  keywords = "polyketide synthase,methyltransferase,oxidoreductase";
                  N = split(keywords, keys, ",")


                  # flag==1 means we are doing regex_match the next N-1 lines
                  # against corresponding array element in keys from [2:N]
                  # once a unmatched found, turn off flag immediately
                  # if the flag==1 reached N-1 lines, then print the good match
                  flag
                  if($NF ~ keys[NR - start_line + 1])
                  F = F ORS $0;
                  if (NR == start_line+N-1) print F > "out_" f++; flag = 0
                  next
                  else
                  flag = 0;



                  # set up the flag/start_line and reset F
                  $NF ~ keys[1] flag = 1; F = $0; start_line= NR;


                  Run the above code with awk -f t24.awk file.txt. You can set up keywords (comma delimited) from your shell(instead of hard-coded in the BEGIN block), and then use -v keywords="..." to make it more flexible.






                  share|improve this answer













                  You might test the following code, where I split your keywords into an awk array named keys with N elements. everything starts with keys[1] and we set up a flag to check the next 1 to N-1 lines if they matches the corresponding values in the array keys [index from 2 to N], any mismatches before the N-1 line will reset this flag, if it reaches the N-1 line, then all are good for output (we also reset flag=0 here so a consecutive run of flag==1 never exceeds N-1 lines):



                  $ cat t24.awk
                  BEGIN
                  FS = OFS = "t";
                  keywords = "polyketide synthase,methyltransferase,oxidoreductase";
                  N = split(keywords, keys, ",")


                  # flag==1 means we are doing regex_match the next N-1 lines
                  # against corresponding array element in keys from [2:N]
                  # once a unmatched found, turn off flag immediately
                  # if the flag==1 reached N-1 lines, then print the good match
                  flag
                  if($NF ~ keys[NR - start_line + 1])
                  F = F ORS $0;
                  if (NR == start_line+N-1) print F > "out_" f++; flag = 0
                  next
                  else
                  flag = 0;



                  # set up the flag/start_line and reset F
                  $NF ~ keys[1] flag = 1; F = $0; start_line= NR;


                  Run the above code with awk -f t24.awk file.txt. You can set up keywords (comma delimited) from your shell(instead of hard-coded in the BEGIN block), and then use -v keywords="..." to make it more flexible.







                  share|improve this answer












                  share|improve this answer



                  share|improve this answer










                  answered yesterday









                  jxcjxc

                  1663




                  1663



























                      draft saved

                      draft discarded
















































                      Thanks for contributing an answer to Unix & Linux Stack Exchange!


                      • Please be sure to answer the question. Provide details and share your research!

                      But avoid


                      • Asking for help, clarification, or responding to other answers.

                      • Making statements based on opinion; back them up with references or personal experience.

                      To learn more, see our tips on writing great answers.




                      draft saved


                      draft discarded














                      StackExchange.ready(
                      function ()
                      StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f511906%2fpick-up-successive-lines-containing-keywords-in-order%23new-answer', 'question_page');

                      );

                      Post as a guest















                      Required, but never shown





















































                      Required, but never shown














                      Required, but never shown












                      Required, but never shown







                      Required, but never shown

































                      Required, but never shown














                      Required, but never shown












                      Required, but never shown







                      Required, but never shown







                      -awk, text-processing

                      Popular posts from this blog

                      Word for a person who has no opinion about whether god existsWord for having a definite opinion while simultaneously withholding judgment?What's the opposite of “newcomer? Is ”veteran" OK?What do you call an “atheist” who might believe in an afterlife?What's a word for someone who wants to voice opinions but not have them challenged?Word for someone who dismisses contrary opinions as irrational?Somone who thinks they are overly special/out of the ordinaryIs there a word, phrase or idiom for “a person who is incapable of thinking about the future”?The belief that a god is human-likeA word for a non-famous person/thing you have heard a lot aboutAdjective for a person who enjoys taking care of their appearance

                      2017 IndyCar Series Contents Series news Teams and drivers Schedule Season summary Footnotes References External links Navigation menu"INDYCAR: Initial 2018 bodywork concepts unveiled"the original"IndyCar confirms switch to Performance Friction brakes in 2017""AJ Foyt Racing will switch to Chevy"the original"Carlos Munoz, Conor Daly will drive for AJ Foyt Racing""Zach Veach's Indy 500 Debut Confirmed with Foyt""No mass exodus from Honda after Ganassi switch""Ex-F1 driver Sato joins Andretti Autosport for 2017 IndyCar season""IndyCar's Ryan Hunter-Reay, sponsor DHL paired through 2020""hhgregg and Andretti Autosport announce partnership for key races in 2016""INDYCAR: Rossi re-signs with Andretti"the original"McLaren Formula 1 - Fernando Alonso to race at Indy 500 with McLaren, Honda and Andretti Autosport""Shank will finally take part in Indy 500 with Harvey, Andretti | MotorSportsTalk""Andretti adds Jack Harvey to Indy 500 field""Ganassi switches to Honda power for 2017""INDYCAR: Chilton returns to Ganassi"the original"IndyCar silly season: Who's going where in 2017?""INDYCAR: Kanaan, NTT Data return to Ganassi"the original"Kimball to remain at Ganassi for 2017""Coyne confirms Bourdais for 2017 IndyCar season""Davison to sub for Bourdais in Indy 500"the original"Gutierrez confirmed for Detroit IndyCar debut""Gutierrez returns with Coyne for rest of 2017 season""Vautier to drive for Coyne at Texas"the original"INDYCAR: Coyne confirms Jones for 2017"the original"Pippa Mann returns to Coyne for Indy 500""Karam, Dreyer & Reinbold teaming up again for Indianapolis 500""Pigot to return to Ed Carpenter Racing""Hildebrand confirmed as full-time Ed Carpenter driver""Veach to replace injured Hildebrand at Barber"the originalNew Team Harding Racing Enters Chaves for 101st Indianapolis 500"Juncos Racing Announces Entry in 101st Running of the Indianapolis 500 :: Juncos Racing""Juncos confirms Pigot for Indy 500""Saavedra confirmed in Juncos' second 500 entry"the original"Lazier confirms Indy 500 run after son's USF2000 debut"the original"Claman DeMelo to race for RLLR at Sonoma"the original"Rahal signs Servia and ace engineer for 2017""IndyCar: Aleshin returns with Schmidt"the original"Aleshin replaced by Saavedra for Toronto""Jack Harvey will pilot SPM No. 7 car at Watkins Glen, Sonoma""Jay Howard confirmed in Tony Stewart's supported SPM Indy entry""INDYCAR: Newgarden to wave the flag at Penske"the original"Pagenaud opts for No. 1 in 2017"the original"Penske confirms Newgarden for 2017""Montoya to stay with Team Penske in 2017""Target leaving IndyCar after 27 seasons with Chip Ganassi""Cavin: IndyCar could see complete driver/team shakeup in 2017""End of the road for KV Racing?""KV Racing confirms closure, equipment sold to Juncos""Juncos confirms IndyCar Series entry"the original"Juncos readies IndyCar program, aims for '17 500"the original"Harding Racing to add Texas, Pocono to schedule"the original"Sato signs with Andretti Autosport for 2017""INDYCAR: Aleshin in Doubt at SPM"the original"Long Beach notebook: JR Hildebrand breaks hand""Hildebrand cleared to return at Phoenix"the original"Bourdais to undergo surgery on multiple fractures""Aleshin loses Schmidt Peterson IndyCar ride""Saavedra in at SPM for Pocono, Gateway"the original"Bourdais to make return at Gateway"the original"The IndyCar Grand Prix no longer is sponsored by Angie's List""2017 IndyCar Series rulebook""2017 Verizon IndyCar Series Official Rulebook"Official websiteeeeee

                      What was this official D&D 3.5e Lovecraft-flavored rulebook?What was this set of RPG tools called?As a first-time DM should I let my players play complex character classes and roles?Nymph's Kiss and the RelationshipWhat was the name of this Cleric Prestige Class that shapes metal with its bare hands?Are the 3.5e Dragonlance books third party or official works?What's up with the domain Vile Darkness?What was this 80s book about RPGs?What was the name of this Werewolf band?What book had Rituals to “upgrade” animal companions to keep them viable at higher levels?What was this RPG that had rules for player-owned businesses?