Using Random Forest variable importance for feature selection Announcing the arrival of Valued Associate #679: Cesar Manara Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern)How does feature selection work in Random Forest?random forests for optimal variable selection/feature selectionIn a random forest algorithm, how can one intrepret the importance of each feature?Random Forest Variable selectionBoruta 'all-relevant' feature selection vs Random Forest 'variables of importance'Can random forest based feature selection method be used for multiple regression in machine learningIn the R randomForest package for random forest feature selection, how is the dataset split for training and testing?Getting feature importance for random forest through cross-validationReducing Bias from a Random Forest - Feature ImportanceInterpreting units for random forest variable importance

Check which numbers satisfy the condition [A*B*C = A! + B! + C!]

Can a USB port passively 'listen only'?

Novel: non-telepath helps overthrow rule by telepaths

When a candle burns, why does the top of wick glow if bottom of flame is hottest?

List of Python versions

Why are Kinder Surprise Eggs illegal in the USA?

What would be the ideal power source for a cybernetic eye?

English words in a non-english sci-fi novel

What does the "x" in "x86" represent?

Okay to merge included columns on otherwise identical indexes?

Understanding Ceva's Theorem

How widely used is the term Treppenwitz? Is it something that most Germans know?

How does the particle を relate to the verb 行く in the structure「A を + B に行く」?

Can an alien society believe that their star system is the universe?

How discoverable are IPv6 addresses and AAAA names by potential attackers?

If a contract sometimes uses the wrong name, is it still valid?

Why light coming from distant stars is not discrete?

List *all* the tuples!

What's the meaning of 間時肆拾貳 at a car parking sign

Why are both D and D# fitting into my E minor key?

How to tell that you are a giant?

Why do people hide their license plates in the EU?

What causes the vertical darker bands in my photo?

What's the purpose of writing one's academic biography in the third person?



Using Random Forest variable importance for feature selection



Announcing the arrival of Valued Associate #679: Cesar Manara
Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern)How does feature selection work in Random Forest?random forests for optimal variable selection/feature selectionIn a random forest algorithm, how can one intrepret the importance of each feature?Random Forest Variable selectionBoruta 'all-relevant' feature selection vs Random Forest 'variables of importance'Can random forest based feature selection method be used for multiple regression in machine learningIn the R randomForest package for random forest feature selection, how is the dataset split for training and testing?Getting feature importance for random forest through cross-validationReducing Bias from a Random Forest - Feature ImportanceInterpreting units for random forest variable importance



.everyoneloves__top-leaderboard:empty,.everyoneloves__mid-leaderboard:empty,.everyoneloves__bot-mid-leaderboard:empty margin-bottom:0;








3












$begingroup$


I'm currently trying to convince my colleague that his method of doing feature selection is causing data leakage and I need help doing so.



The method they are using is as follows:
They first run a random forest on all variables and get the feature importance measure; MeanDecreaseAccuracy. They then remove all variables that score low on this measure and re-run the forest and report the out of bag error rate as the error for the model.



They argue that since the MeanDecreaseAccuracy measure is calculated using the bootstrap and out of bag records that there is no data leakage. I am trying to convince them that since the variable importance measure uses ALL data (in bag records to build the trees and out of bag records to obtain the decrease in accuracy) there is data leakage if they use this measure to do feature selection in this manner.



My solution for them was that they cannot use the out of bag error measure if they want to do feature selection, they will have to set up a proper cross validation split and perform the feature selection on the training sets only.



Am I incorrect here? Can anyone think of a convincing argument (example or paper) that I can show my colleague?










share|cite|improve this question









$endgroup$











  • $begingroup$
    you might like to read up on boruta
    $endgroup$
    – Sycorax
    15 hours ago










  • $begingroup$
    Yes I am familiar with boruta, but this isn't about the merits of boruta vs. traditional random forest importance, rather the correct implication of feature selection. From what I understand of boruta it would have the same issues if applied in the manner I describe above.
    $endgroup$
    – astel
    12 hours ago










  • $begingroup$
    sure, they're a way to misuse any tool. but I've found it helpful to "yes-but" the bad ideas that other people have instead of saying "no", as in "Yes, we can use random forest to do feature selection. Here's a good way that we can do that..."
    $endgroup$
    – Sycorax
    12 hours ago

















3












$begingroup$


I'm currently trying to convince my colleague that his method of doing feature selection is causing data leakage and I need help doing so.



The method they are using is as follows:
They first run a random forest on all variables and get the feature importance measure; MeanDecreaseAccuracy. They then remove all variables that score low on this measure and re-run the forest and report the out of bag error rate as the error for the model.



They argue that since the MeanDecreaseAccuracy measure is calculated using the bootstrap and out of bag records that there is no data leakage. I am trying to convince them that since the variable importance measure uses ALL data (in bag records to build the trees and out of bag records to obtain the decrease in accuracy) there is data leakage if they use this measure to do feature selection in this manner.



My solution for them was that they cannot use the out of bag error measure if they want to do feature selection, they will have to set up a proper cross validation split and perform the feature selection on the training sets only.



Am I incorrect here? Can anyone think of a convincing argument (example or paper) that I can show my colleague?










share|cite|improve this question









$endgroup$











  • $begingroup$
    you might like to read up on boruta
    $endgroup$
    – Sycorax
    15 hours ago










  • $begingroup$
    Yes I am familiar with boruta, but this isn't about the merits of boruta vs. traditional random forest importance, rather the correct implication of feature selection. From what I understand of boruta it would have the same issues if applied in the manner I describe above.
    $endgroup$
    – astel
    12 hours ago










  • $begingroup$
    sure, they're a way to misuse any tool. but I've found it helpful to "yes-but" the bad ideas that other people have instead of saying "no", as in "Yes, we can use random forest to do feature selection. Here's a good way that we can do that..."
    $endgroup$
    – Sycorax
    12 hours ago













3












3








3





$begingroup$


I'm currently trying to convince my colleague that his method of doing feature selection is causing data leakage and I need help doing so.



The method they are using is as follows:
They first run a random forest on all variables and get the feature importance measure; MeanDecreaseAccuracy. They then remove all variables that score low on this measure and re-run the forest and report the out of bag error rate as the error for the model.



They argue that since the MeanDecreaseAccuracy measure is calculated using the bootstrap and out of bag records that there is no data leakage. I am trying to convince them that since the variable importance measure uses ALL data (in bag records to build the trees and out of bag records to obtain the decrease in accuracy) there is data leakage if they use this measure to do feature selection in this manner.



My solution for them was that they cannot use the out of bag error measure if they want to do feature selection, they will have to set up a proper cross validation split and perform the feature selection on the training sets only.



Am I incorrect here? Can anyone think of a convincing argument (example or paper) that I can show my colleague?










share|cite|improve this question









$endgroup$




I'm currently trying to convince my colleague that his method of doing feature selection is causing data leakage and I need help doing so.



The method they are using is as follows:
They first run a random forest on all variables and get the feature importance measure; MeanDecreaseAccuracy. They then remove all variables that score low on this measure and re-run the forest and report the out of bag error rate as the error for the model.



They argue that since the MeanDecreaseAccuracy measure is calculated using the bootstrap and out of bag records that there is no data leakage. I am trying to convince them that since the variable importance measure uses ALL data (in bag records to build the trees and out of bag records to obtain the decrease in accuracy) there is data leakage if they use this measure to do feature selection in this manner.



My solution for them was that they cannot use the out of bag error measure if they want to do feature selection, they will have to set up a proper cross validation split and perform the feature selection on the training sets only.



Am I incorrect here? Can anyone think of a convincing argument (example or paper) that I can show my colleague?







feature-selection random-forest bootstrap data-leakage






share|cite|improve this question













share|cite|improve this question











share|cite|improve this question




share|cite|improve this question










asked 15 hours ago









astelastel

311113




311113











  • $begingroup$
    you might like to read up on boruta
    $endgroup$
    – Sycorax
    15 hours ago










  • $begingroup$
    Yes I am familiar with boruta, but this isn't about the merits of boruta vs. traditional random forest importance, rather the correct implication of feature selection. From what I understand of boruta it would have the same issues if applied in the manner I describe above.
    $endgroup$
    – astel
    12 hours ago










  • $begingroup$
    sure, they're a way to misuse any tool. but I've found it helpful to "yes-but" the bad ideas that other people have instead of saying "no", as in "Yes, we can use random forest to do feature selection. Here's a good way that we can do that..."
    $endgroup$
    – Sycorax
    12 hours ago
















  • $begingroup$
    you might like to read up on boruta
    $endgroup$
    – Sycorax
    15 hours ago










  • $begingroup$
    Yes I am familiar with boruta, but this isn't about the merits of boruta vs. traditional random forest importance, rather the correct implication of feature selection. From what I understand of boruta it would have the same issues if applied in the manner I describe above.
    $endgroup$
    – astel
    12 hours ago










  • $begingroup$
    sure, they're a way to misuse any tool. but I've found it helpful to "yes-but" the bad ideas that other people have instead of saying "no", as in "Yes, we can use random forest to do feature selection. Here's a good way that we can do that..."
    $endgroup$
    – Sycorax
    12 hours ago















$begingroup$
you might like to read up on boruta
$endgroup$
– Sycorax
15 hours ago




$begingroup$
you might like to read up on boruta
$endgroup$
– Sycorax
15 hours ago












$begingroup$
Yes I am familiar with boruta, but this isn't about the merits of boruta vs. traditional random forest importance, rather the correct implication of feature selection. From what I understand of boruta it would have the same issues if applied in the manner I describe above.
$endgroup$
– astel
12 hours ago




$begingroup$
Yes I am familiar with boruta, but this isn't about the merits of boruta vs. traditional random forest importance, rather the correct implication of feature selection. From what I understand of boruta it would have the same issues if applied in the manner I describe above.
$endgroup$
– astel
12 hours ago












$begingroup$
sure, they're a way to misuse any tool. but I've found it helpful to "yes-but" the bad ideas that other people have instead of saying "no", as in "Yes, we can use random forest to do feature selection. Here's a good way that we can do that..."
$endgroup$
– Sycorax
12 hours ago




$begingroup$
sure, they're a way to misuse any tool. but I've found it helpful to "yes-but" the bad ideas that other people have instead of saying "no", as in "Yes, we can use random forest to do feature selection. Here's a good way that we can do that..."
$endgroup$
– Sycorax
12 hours ago










1 Answer
1






active

oldest

votes


















4












$begingroup$

You are entirely correct!



A couple of months ago I was in the exact same position when justifying a different feature selection approach in front of my supervisors. I will cite the sentence I used in my thesis, although it has not been published yet.




Since the ordering of the variables depends on all samples, the
selection step is performed using information of all samples and
thus, the OOB error of the subsequent model no longer has the
properties of an independent test set as it is not independent from
the previous selection step.



- Marc H.




For references, see section 4.1 of 'A new variable selection approach using Random Forests' by Hapfelmeier and Ulm or 'Application of Breiman’s Random Forest to Modeling Structure-Activity Relationships of Pharmaceutical Molecules
' by Svetnik et al., who address this issue in context of forward-/backward feature selection.






share|cite|improve this answer









$endgroup$













    Your Answer








    StackExchange.ready(function()
    var channelOptions =
    tags: "".split(" "),
    id: "65"
    ;
    initTagRenderer("".split(" "), "".split(" "), channelOptions);

    StackExchange.using("externalEditor", function()
    // Have to fire editor after snippets, if snippets enabled
    if (StackExchange.settings.snippets.snippetsEnabled)
    StackExchange.using("snippets", function()
    createEditor();
    );

    else
    createEditor();

    );

    function createEditor()
    StackExchange.prepareEditor(
    heartbeatType: 'answer',
    autoActivateHeartbeat: false,
    convertImagesToLinks: false,
    noModals: true,
    showLowRepImageUploadWarning: true,
    reputationToPostImages: null,
    bindNavPrevention: true,
    postfix: "",
    imageUploader:
    brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
    contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
    allowUrls: true
    ,
    onDemand: true,
    discardSelector: ".discard-answer"
    ,immediatelyShowMarkdownHelp:true
    );



    );













    draft saved

    draft discarded


















    StackExchange.ready(
    function ()
    StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fstats.stackexchange.com%2fquestions%2f403381%2fusing-random-forest-variable-importance-for-feature-selection%23new-answer', 'question_page');

    );

    Post as a guest















    Required, but never shown

























    1 Answer
    1






    active

    oldest

    votes








    1 Answer
    1






    active

    oldest

    votes









    active

    oldest

    votes






    active

    oldest

    votes









    4












    $begingroup$

    You are entirely correct!



    A couple of months ago I was in the exact same position when justifying a different feature selection approach in front of my supervisors. I will cite the sentence I used in my thesis, although it has not been published yet.




    Since the ordering of the variables depends on all samples, the
    selection step is performed using information of all samples and
    thus, the OOB error of the subsequent model no longer has the
    properties of an independent test set as it is not independent from
    the previous selection step.



    - Marc H.




    For references, see section 4.1 of 'A new variable selection approach using Random Forests' by Hapfelmeier and Ulm or 'Application of Breiman’s Random Forest to Modeling Structure-Activity Relationships of Pharmaceutical Molecules
    ' by Svetnik et al., who address this issue in context of forward-/backward feature selection.






    share|cite|improve this answer









    $endgroup$

















      4












      $begingroup$

      You are entirely correct!



      A couple of months ago I was in the exact same position when justifying a different feature selection approach in front of my supervisors. I will cite the sentence I used in my thesis, although it has not been published yet.




      Since the ordering of the variables depends on all samples, the
      selection step is performed using information of all samples and
      thus, the OOB error of the subsequent model no longer has the
      properties of an independent test set as it is not independent from
      the previous selection step.



      - Marc H.




      For references, see section 4.1 of 'A new variable selection approach using Random Forests' by Hapfelmeier and Ulm or 'Application of Breiman’s Random Forest to Modeling Structure-Activity Relationships of Pharmaceutical Molecules
      ' by Svetnik et al., who address this issue in context of forward-/backward feature selection.






      share|cite|improve this answer









      $endgroup$















        4












        4








        4





        $begingroup$

        You are entirely correct!



        A couple of months ago I was in the exact same position when justifying a different feature selection approach in front of my supervisors. I will cite the sentence I used in my thesis, although it has not been published yet.




        Since the ordering of the variables depends on all samples, the
        selection step is performed using information of all samples and
        thus, the OOB error of the subsequent model no longer has the
        properties of an independent test set as it is not independent from
        the previous selection step.



        - Marc H.




        For references, see section 4.1 of 'A new variable selection approach using Random Forests' by Hapfelmeier and Ulm or 'Application of Breiman’s Random Forest to Modeling Structure-Activity Relationships of Pharmaceutical Molecules
        ' by Svetnik et al., who address this issue in context of forward-/backward feature selection.






        share|cite|improve this answer









        $endgroup$



        You are entirely correct!



        A couple of months ago I was in the exact same position when justifying a different feature selection approach in front of my supervisors. I will cite the sentence I used in my thesis, although it has not been published yet.




        Since the ordering of the variables depends on all samples, the
        selection step is performed using information of all samples and
        thus, the OOB error of the subsequent model no longer has the
        properties of an independent test set as it is not independent from
        the previous selection step.



        - Marc H.




        For references, see section 4.1 of 'A new variable selection approach using Random Forests' by Hapfelmeier and Ulm or 'Application of Breiman’s Random Forest to Modeling Structure-Activity Relationships of Pharmaceutical Molecules
        ' by Svetnik et al., who address this issue in context of forward-/backward feature selection.







        share|cite|improve this answer












        share|cite|improve this answer



        share|cite|improve this answer










        answered 14 hours ago









        bi_scholarbi_scholar

        50113




        50113



























            draft saved

            draft discarded
















































            Thanks for contributing an answer to Cross Validated!


            • Please be sure to answer the question. Provide details and share your research!

            But avoid


            • Asking for help, clarification, or responding to other answers.

            • Making statements based on opinion; back them up with references or personal experience.

            Use MathJax to format equations. MathJax reference.


            To learn more, see our tips on writing great answers.




            draft saved


            draft discarded














            StackExchange.ready(
            function ()
            StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fstats.stackexchange.com%2fquestions%2f403381%2fusing-random-forest-variable-importance-for-feature-selection%23new-answer', 'question_page');

            );

            Post as a guest















            Required, but never shown





















































            Required, but never shown














            Required, but never shown












            Required, but never shown







            Required, but never shown

































            Required, but never shown














            Required, but never shown












            Required, but never shown







            Required, but never shown







            -bootstrap, data-leakage, feature-selection, random-forest

            Popular posts from this blog

            Word for a person who has no opinion about whether god existsWord for having a definite opinion while simultaneously withholding judgment?What's the opposite of “newcomer? Is ”veteran" OK?What do you call an “atheist” who might believe in an afterlife?What's a word for someone who wants to voice opinions but not have them challenged?Word for someone who dismisses contrary opinions as irrational?Somone who thinks they are overly special/out of the ordinaryIs there a word, phrase or idiom for “a person who is incapable of thinking about the future”?The belief that a god is human-likeA word for a non-famous person/thing you have heard a lot aboutAdjective for a person who enjoys taking care of their appearance

            What was this official D&D 3.5e Lovecraft-flavored rulebook?What was this set of RPG tools called?As a first-time DM should I let my players play complex character classes and roles?Nymph's Kiss and the RelationshipWhat was the name of this Cleric Prestige Class that shapes metal with its bare hands?Are the 3.5e Dragonlance books third party or official works?What's up with the domain Vile Darkness?What was this 80s book about RPGs?What was the name of this Werewolf band?What book had Rituals to “upgrade” animal companions to keep them viable at higher levels?What was this RPG that had rules for player-owned businesses?

            2017 IndyCar Series Contents Series news Teams and drivers Schedule Season summary Footnotes References External links Navigation menu"INDYCAR: Initial 2018 bodywork concepts unveiled"the original"IndyCar confirms switch to Performance Friction brakes in 2017""AJ Foyt Racing will switch to Chevy"the original"Carlos Munoz, Conor Daly will drive for AJ Foyt Racing""Zach Veach's Indy 500 Debut Confirmed with Foyt""No mass exodus from Honda after Ganassi switch""Ex-F1 driver Sato joins Andretti Autosport for 2017 IndyCar season""IndyCar's Ryan Hunter-Reay, sponsor DHL paired through 2020""hhgregg and Andretti Autosport announce partnership for key races in 2016""INDYCAR: Rossi re-signs with Andretti"the original"McLaren Formula 1 - Fernando Alonso to race at Indy 500 with McLaren, Honda and Andretti Autosport""Shank will finally take part in Indy 500 with Harvey, Andretti | MotorSportsTalk""Andretti adds Jack Harvey to Indy 500 field""Ganassi switches to Honda power for 2017""INDYCAR: Chilton returns to Ganassi"the original"IndyCar silly season: Who's going where in 2017?""INDYCAR: Kanaan, NTT Data return to Ganassi"the original"Kimball to remain at Ganassi for 2017""Coyne confirms Bourdais for 2017 IndyCar season""Davison to sub for Bourdais in Indy 500"the original"Gutierrez confirmed for Detroit IndyCar debut""Gutierrez returns with Coyne for rest of 2017 season""Vautier to drive for Coyne at Texas"the original"INDYCAR: Coyne confirms Jones for 2017"the original"Pippa Mann returns to Coyne for Indy 500""Karam, Dreyer & Reinbold teaming up again for Indianapolis 500""Pigot to return to Ed Carpenter Racing""Hildebrand confirmed as full-time Ed Carpenter driver""Veach to replace injured Hildebrand at Barber"the originalNew Team Harding Racing Enters Chaves for 101st Indianapolis 500"Juncos Racing Announces Entry in 101st Running of the Indianapolis 500 :: Juncos Racing""Juncos confirms Pigot for Indy 500""Saavedra confirmed in Juncos' second 500 entry"the original"Lazier confirms Indy 500 run after son's USF2000 debut"the original"Claman DeMelo to race for RLLR at Sonoma"the original"Rahal signs Servia and ace engineer for 2017""IndyCar: Aleshin returns with Schmidt"the original"Aleshin replaced by Saavedra for Toronto""Jack Harvey will pilot SPM No. 7 car at Watkins Glen, Sonoma""Jay Howard confirmed in Tony Stewart's supported SPM Indy entry""INDYCAR: Newgarden to wave the flag at Penske"the original"Pagenaud opts for No. 1 in 2017"the original"Penske confirms Newgarden for 2017""Montoya to stay with Team Penske in 2017""Target leaving IndyCar after 27 seasons with Chip Ganassi""Cavin: IndyCar could see complete driver/team shakeup in 2017""End of the road for KV Racing?""KV Racing confirms closure, equipment sold to Juncos""Juncos confirms IndyCar Series entry"the original"Juncos readies IndyCar program, aims for '17 500"the original"Harding Racing to add Texas, Pocono to schedule"the original"Sato signs with Andretti Autosport for 2017""INDYCAR: Aleshin in Doubt at SPM"the original"Long Beach notebook: JR Hildebrand breaks hand""Hildebrand cleared to return at Phoenix"the original"Bourdais to undergo surgery on multiple fractures""Aleshin loses Schmidt Peterson IndyCar ride""Saavedra in at SPM for Pocono, Gateway"the original"Bourdais to make return at Gateway"the original"The IndyCar Grand Prix no longer is sponsored by Angie's List""2017 IndyCar Series rulebook""2017 Verizon IndyCar Series Official Rulebook"Official websiteeeeee