Extracting lines to new files The 2019 Stack Overflow Developer Survey Results Are In Announcing the arrival of Valued Associate #679: Cesar Manara Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern) 2019 Community Moderator Election ResultsExtracting file name and string from multiple filesExtracting part of lines with specific pattern using awk,sedhow to print new word after two lines using awkExtracting lines based on conditionsextracting date field from the linesinsert new lines into a csv file obtained via curl on an apiarithmetic operations within column with awk or sedExtracting pattern from multiple linesprint out lines if first three columns match the first three columns in another fileSplitting text file into CSV with multiple delimiters in bash?

Do ℕ, mathbbN, BbbN, symbbN effectively differ, and is there a "canonical" specification of the naturals?

Why doesn't shell automatically fix "useless use of cat"?

How many cones with angle theta can I pack into the unit sphere?

Is there a way to generate uniformly distributed points on a sphere from a fixed amount of random real numbers per point?

What information about me do stores get via my credit card?

Accepted by European university, rejected by all American ones I applied to? Possible reasons?

Button changing its text & action. Good or terrible?

Is 'stolen' appropriate word?

Python - Fishing Simulator

Are spiders unable to hurt humans, especially very small spiders?

Word to describe a time interval

Deal with toxic manager when you can't quit

Does Parliament need to approve the new Brexit delay to 31 October 2019?

How did the audience guess the pentatonic scale in Bobby McFerrin's presentation?

Identify 80s or 90s comics with ripped creatures (not dwarves)

Is there a writing software that you can sort scenes like slides in PowerPoint?

Circular reasoning in L'Hopital's rule

What was the last x86 CPU that did not have the x87 floating-point unit built in?

"is" operation returns false even though two objects have same id

The following signatures were invalid: EXPKEYSIG 1397BC53640DB551

should truth entail possible truth

How to politely respond to generic emails requesting a PhD/job in my lab? Without wasting too much time

Using dividends to reduce short term capital gains?

Match Roman Numerals



Extracting lines to new files



The 2019 Stack Overflow Developer Survey Results Are In
Announcing the arrival of Valued Associate #679: Cesar Manara
Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern)
2019 Community Moderator Election ResultsExtracting file name and string from multiple filesExtracting part of lines with specific pattern using awk,sedhow to print new word after two lines using awkExtracting lines based on conditionsextracting date field from the linesinsert new lines into a csv file obtained via curl on an apiarithmetic operations within column with awk or sedExtracting pattern from multiple linesprint out lines if first three columns match the first three columns in another fileSplitting text file into CSV with multiple delimiters in bash?



.everyoneloves__top-leaderboard:empty,.everyoneloves__mid-leaderboard:empty,.everyoneloves__bot-mid-leaderboard:empty margin-bottom:0;








0















Say I have a large CSV file with a header and several columns. For the purpose of this question I will consider a small file with just two columns. We can call it use_rep.



user_id,rep
885,500K+
22565,200K+
7453,200K+
86440,100K+
116858,100K+
22222,100K+
38906,100K+
10762,<100K
70524,<100K


I'd like to send each row to a file corresponding to the value on the second column. For example, I'd like there to be a file whose name is 200K+ and whose content is



user_id,rep
22565,200K+
7453,200K+


The contents of use_rep should not be assumed to be ordered in anyway. The pattern to be used would ideally accept regular expressions.



No sed or perl is preferred.










share|improve this question









New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.




















  • I think AWK can do this easily, but I don't really know how.

    – regex
    yesterday

















0















Say I have a large CSV file with a header and several columns. For the purpose of this question I will consider a small file with just two columns. We can call it use_rep.



user_id,rep
885,500K+
22565,200K+
7453,200K+
86440,100K+
116858,100K+
22222,100K+
38906,100K+
10762,<100K
70524,<100K


I'd like to send each row to a file corresponding to the value on the second column. For example, I'd like there to be a file whose name is 200K+ and whose content is



user_id,rep
22565,200K+
7453,200K+


The contents of use_rep should not be assumed to be ordered in anyway. The pattern to be used would ideally accept regular expressions.



No sed or perl is preferred.










share|improve this question









New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.




















  • I think AWK can do this easily, but I don't really know how.

    – regex
    yesterday













0












0








0








Say I have a large CSV file with a header and several columns. For the purpose of this question I will consider a small file with just two columns. We can call it use_rep.



user_id,rep
885,500K+
22565,200K+
7453,200K+
86440,100K+
116858,100K+
22222,100K+
38906,100K+
10762,<100K
70524,<100K


I'd like to send each row to a file corresponding to the value on the second column. For example, I'd like there to be a file whose name is 200K+ and whose content is



user_id,rep
22565,200K+
7453,200K+


The contents of use_rep should not be assumed to be ordered in anyway. The pattern to be used would ideally accept regular expressions.



No sed or perl is preferred.










share|improve this question









New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.












Say I have a large CSV file with a header and several columns. For the purpose of this question I will consider a small file with just two columns. We can call it use_rep.



user_id,rep
885,500K+
22565,200K+
7453,200K+
86440,100K+
116858,100K+
22222,100K+
38906,100K+
10762,<100K
70524,<100K


I'd like to send each row to a file corresponding to the value on the second column. For example, I'd like there to be a file whose name is 200K+ and whose content is



user_id,rep
22565,200K+
7453,200K+


The contents of use_rep should not be assumed to be ordered in anyway. The pattern to be used would ideally accept regular expressions.



No sed or perl is preferred.







text-processing awk pattern-matching






share|improve this question









New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.











share|improve this question









New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.









share|improve this question




share|improve this question








edited yesterday







regex













New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.









asked yesterday









regexregex

223




223




New contributor




regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.





New contributor





regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.






regex is a new contributor to this site. Take care in asking for clarification, commenting, and answering.
Check out our Code of Conduct.












  • I think AWK can do this easily, but I don't really know how.

    – regex
    yesterday

















  • I think AWK can do this easily, but I don't really know how.

    – regex
    yesterday
















I think AWK can do this easily, but I don't really know how.

– regex
yesterday





I think AWK can do this easily, but I don't really know how.

– regex
yesterday










1 Answer
1






active

oldest

votes


















3














Ignoring the header (which you can tack on later):



awk -F, 'NR > 1 print > $2' use_rep


which will print each line to a file named by the second column:



~ head *[0-9]*
==> 100K+ <==
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
22565,200K+
7453,200K+

==> 500K+ <==
885,500K+

==> <100K <==
10762,<100K


To put the header, maybe something like:



awk -F, 'NR == 1 header = $0; next # save header, skip this line
!a[$2]++ print header > $2 # print if second field hasnt been seen before
print > $2 ' use_rep


Result:



~ head *[0-9]*
==> 100K+ <==
user_id,rep
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
user_id,rep
22565,200K+
7453,200K+

==> 500K+ <==
user_id,rep
885,500K+

==> <100K <==
user_id,rep
10762,<100K
70524,<100K





share|improve this answer

























  • I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

    – regex
    yesterday











  • Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

    – muru
    yesterday











Your Answer








StackExchange.ready(function()
var channelOptions =
tags: "".split(" "),
id: "106"
;
initTagRenderer("".split(" "), "".split(" "), channelOptions);

StackExchange.using("externalEditor", function()
// Have to fire editor after snippets, if snippets enabled
if (StackExchange.settings.snippets.snippetsEnabled)
StackExchange.using("snippets", function()
createEditor();
);

else
createEditor();

);

function createEditor()
StackExchange.prepareEditor(
heartbeatType: 'answer',
autoActivateHeartbeat: false,
convertImagesToLinks: false,
noModals: true,
showLowRepImageUploadWarning: true,
reputationToPostImages: null,
bindNavPrevention: true,
postfix: "",
imageUploader:
brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
allowUrls: true
,
onDemand: true,
discardSelector: ".discard-answer"
,immediatelyShowMarkdownHelp:true
);



);






regex is a new contributor. Be nice, and check out our Code of Conduct.









draft saved

draft discarded


















StackExchange.ready(
function ()
StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f511886%2fextracting-lines-to-new-files%23new-answer', 'question_page');

);

Post as a guest















Required, but never shown

























1 Answer
1






active

oldest

votes








1 Answer
1






active

oldest

votes









active

oldest

votes






active

oldest

votes









3














Ignoring the header (which you can tack on later):



awk -F, 'NR > 1 print > $2' use_rep


which will print each line to a file named by the second column:



~ head *[0-9]*
==> 100K+ <==
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
22565,200K+
7453,200K+

==> 500K+ <==
885,500K+

==> <100K <==
10762,<100K


To put the header, maybe something like:



awk -F, 'NR == 1 header = $0; next # save header, skip this line
!a[$2]++ print header > $2 # print if second field hasnt been seen before
print > $2 ' use_rep


Result:



~ head *[0-9]*
==> 100K+ <==
user_id,rep
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
user_id,rep
22565,200K+
7453,200K+

==> 500K+ <==
user_id,rep
885,500K+

==> <100K <==
user_id,rep
10762,<100K
70524,<100K





share|improve this answer

























  • I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

    – regex
    yesterday











  • Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

    – muru
    yesterday















3














Ignoring the header (which you can tack on later):



awk -F, 'NR > 1 print > $2' use_rep


which will print each line to a file named by the second column:



~ head *[0-9]*
==> 100K+ <==
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
22565,200K+
7453,200K+

==> 500K+ <==
885,500K+

==> <100K <==
10762,<100K


To put the header, maybe something like:



awk -F, 'NR == 1 header = $0; next # save header, skip this line
!a[$2]++ print header > $2 # print if second field hasnt been seen before
print > $2 ' use_rep


Result:



~ head *[0-9]*
==> 100K+ <==
user_id,rep
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
user_id,rep
22565,200K+
7453,200K+

==> 500K+ <==
user_id,rep
885,500K+

==> <100K <==
user_id,rep
10762,<100K
70524,<100K





share|improve this answer

























  • I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

    – regex
    yesterday











  • Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

    – muru
    yesterday













3












3








3







Ignoring the header (which you can tack on later):



awk -F, 'NR > 1 print > $2' use_rep


which will print each line to a file named by the second column:



~ head *[0-9]*
==> 100K+ <==
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
22565,200K+
7453,200K+

==> 500K+ <==
885,500K+

==> <100K <==
10762,<100K


To put the header, maybe something like:



awk -F, 'NR == 1 header = $0; next # save header, skip this line
!a[$2]++ print header > $2 # print if second field hasnt been seen before
print > $2 ' use_rep


Result:



~ head *[0-9]*
==> 100K+ <==
user_id,rep
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
user_id,rep
22565,200K+
7453,200K+

==> 500K+ <==
user_id,rep
885,500K+

==> <100K <==
user_id,rep
10762,<100K
70524,<100K





share|improve this answer















Ignoring the header (which you can tack on later):



awk -F, 'NR > 1 print > $2' use_rep


which will print each line to a file named by the second column:



~ head *[0-9]*
==> 100K+ <==
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
22565,200K+
7453,200K+

==> 500K+ <==
885,500K+

==> <100K <==
10762,<100K


To put the header, maybe something like:



awk -F, 'NR == 1 header = $0; next # save header, skip this line
!a[$2]++ print header > $2 # print if second field hasnt been seen before
print > $2 ' use_rep


Result:



~ head *[0-9]*
==> 100K+ <==
user_id,rep
86440,100K+
116858,100K+
22222,100K+
38906,100K+

==> 200K+ <==
user_id,rep
22565,200K+
7453,200K+

==> 500K+ <==
user_id,rep
885,500K+

==> <100K <==
user_id,rep
10762,<100K
70524,<100K






share|improve this answer














share|improve this answer



share|improve this answer








edited yesterday









regex

223




223










answered yesterday









murumuru

37.6k589165




37.6k589165












  • I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

    – regex
    yesterday











  • Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

    – muru
    yesterday

















  • I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

    – regex
    yesterday











  • Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

    – muru
    yesterday
















I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

– regex
yesterday





I'm having some issues with this due to commas inside text identifiers (" "). Is there an easy fix?

– regex
yesterday













Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

– muru
yesterday





Not with awk. You should use a tool with support for quoted csv, like csvkit or Python or Perl.

– muru
yesterday










regex is a new contributor. Be nice, and check out our Code of Conduct.









draft saved

draft discarded


















regex is a new contributor. Be nice, and check out our Code of Conduct.












regex is a new contributor. Be nice, and check out our Code of Conduct.











regex is a new contributor. Be nice, and check out our Code of Conduct.














Thanks for contributing an answer to Unix & Linux Stack Exchange!


  • Please be sure to answer the question. Provide details and share your research!

But avoid


  • Asking for help, clarification, or responding to other answers.

  • Making statements based on opinion; back them up with references or personal experience.

To learn more, see our tips on writing great answers.




draft saved


draft discarded














StackExchange.ready(
function ()
StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f511886%2fextracting-lines-to-new-files%23new-answer', 'question_page');

);

Post as a guest















Required, but never shown





















































Required, but never shown














Required, but never shown












Required, but never shown







Required, but never shown

































Required, but never shown














Required, but never shown












Required, but never shown







Required, but never shown







-awk, pattern-matching, text-processing

Popular posts from this blog

Word for a person who has no opinion about whether god existsWord for having a definite opinion while simultaneously withholding judgment?What's the opposite of “newcomer? Is ”veteran" OK?What do you call an “atheist” who might believe in an afterlife?What's a word for someone who wants to voice opinions but not have them challenged?Word for someone who dismisses contrary opinions as irrational?Somone who thinks they are overly special/out of the ordinaryIs there a word, phrase or idiom for “a person who is incapable of thinking about the future”?The belief that a god is human-likeA word for a non-famous person/thing you have heard a lot aboutAdjective for a person who enjoys taking care of their appearance

What was this official D&D 3.5e Lovecraft-flavored rulebook?What was this set of RPG tools called?As a first-time DM should I let my players play complex character classes and roles?Nymph's Kiss and the RelationshipWhat was the name of this Cleric Prestige Class that shapes metal with its bare hands?Are the 3.5e Dragonlance books third party or official works?What's up with the domain Vile Darkness?What was this 80s book about RPGs?What was the name of this Werewolf band?What book had Rituals to “upgrade” animal companions to keep them viable at higher levels?What was this RPG that had rules for player-owned businesses?

2017 IndyCar Series Contents Series news Teams and drivers Schedule Season summary Footnotes References External links Navigation menu"INDYCAR: Initial 2018 bodywork concepts unveiled"the original"IndyCar confirms switch to Performance Friction brakes in 2017""AJ Foyt Racing will switch to Chevy"the original"Carlos Munoz, Conor Daly will drive for AJ Foyt Racing""Zach Veach's Indy 500 Debut Confirmed with Foyt""No mass exodus from Honda after Ganassi switch""Ex-F1 driver Sato joins Andretti Autosport for 2017 IndyCar season""IndyCar's Ryan Hunter-Reay, sponsor DHL paired through 2020""hhgregg and Andretti Autosport announce partnership for key races in 2016""INDYCAR: Rossi re-signs with Andretti"the original"McLaren Formula 1 - Fernando Alonso to race at Indy 500 with McLaren, Honda and Andretti Autosport""Shank will finally take part in Indy 500 with Harvey, Andretti | MotorSportsTalk""Andretti adds Jack Harvey to Indy 500 field""Ganassi switches to Honda power for 2017""INDYCAR: Chilton returns to Ganassi"the original"IndyCar silly season: Who's going where in 2017?""INDYCAR: Kanaan, NTT Data return to Ganassi"the original"Kimball to remain at Ganassi for 2017""Coyne confirms Bourdais for 2017 IndyCar season""Davison to sub for Bourdais in Indy 500"the original"Gutierrez confirmed for Detroit IndyCar debut""Gutierrez returns with Coyne for rest of 2017 season""Vautier to drive for Coyne at Texas"the original"INDYCAR: Coyne confirms Jones for 2017"the original"Pippa Mann returns to Coyne for Indy 500""Karam, Dreyer & Reinbold teaming up again for Indianapolis 500""Pigot to return to Ed Carpenter Racing""Hildebrand confirmed as full-time Ed Carpenter driver""Veach to replace injured Hildebrand at Barber"the originalNew Team Harding Racing Enters Chaves for 101st Indianapolis 500"Juncos Racing Announces Entry in 101st Running of the Indianapolis 500 :: Juncos Racing""Juncos confirms Pigot for Indy 500""Saavedra confirmed in Juncos' second 500 entry"the original"Lazier confirms Indy 500 run after son's USF2000 debut"the original"Claman DeMelo to race for RLLR at Sonoma"the original"Rahal signs Servia and ace engineer for 2017""IndyCar: Aleshin returns with Schmidt"the original"Aleshin replaced by Saavedra for Toronto""Jack Harvey will pilot SPM No. 7 car at Watkins Glen, Sonoma""Jay Howard confirmed in Tony Stewart's supported SPM Indy entry""INDYCAR: Newgarden to wave the flag at Penske"the original"Pagenaud opts for No. 1 in 2017"the original"Penske confirms Newgarden for 2017""Montoya to stay with Team Penske in 2017""Target leaving IndyCar after 27 seasons with Chip Ganassi""Cavin: IndyCar could see complete driver/team shakeup in 2017""End of the road for KV Racing?""KV Racing confirms closure, equipment sold to Juncos""Juncos confirms IndyCar Series entry"the original"Juncos readies IndyCar program, aims for '17 500"the original"Harding Racing to add Texas, Pocono to schedule"the original"Sato signs with Andretti Autosport for 2017""INDYCAR: Aleshin in Doubt at SPM"the original"Long Beach notebook: JR Hildebrand breaks hand""Hildebrand cleared to return at Phoenix"the original"Bourdais to undergo surgery on multiple fractures""Aleshin loses Schmidt Peterson IndyCar ride""Saavedra in at SPM for Pocono, Gateway"the original"Bourdais to make return at Gateway"the original"The IndyCar Grand Prix no longer is sponsored by Angie's List""2017 IndyCar Series rulebook""2017 Verizon IndyCar Series Official Rulebook"Official websiteeeeee