Staskoverflow! I am a complete noob in Machine Learning in Python. I did go through a bunch of books and practice a ton of code. It seems like for more or less intelligent model selection one needs to do some version of gridsearch (I'm looking into Randomized CV Search as the most optimal thing). So I'm making my pipelines - at this stage I wrote a custom column selector by dtype and moved towards custom imputers and standardizers. After transformations I'm successfully doing a feature union for further fun.
My column selector returns a dataframe with all original column names. The imputer that follows it in the pipeline returns a numpy array. All the following transformers like a standardizer obviously get an array and return an array. I examined the imputer with dir() and got this output:
['__class__', '__delattr__', '__dict__', '__dir__', '__doc__', '__eq__',
'__format__', '__ge__', '__getattribute__', '__getstate__', '__gt__',
'__hash__', '__init__', '__init_subclass__', '__le__', '__lt__',
'__module__', '__ne__', '__new__', '__reduce__', '__reduce_ex__',
'__repr__', '__setattr__', '__setstate__', '__sizeof__', '__str__',
'__subclasshook__', '__weakref__', '_dense_fit', '_get_param_names',
'_sparse_fit', 'axis', 'copy', 'fit', 'fit_transform', 'get_params',
'missing_values', 'set_params', 'statistics_', 'strategy', 'transform',
'verbose']
I examined everything and nowhere does the imputer store original data or its column names.
Does this mean that I need to either directly modify each of such sklearn classes (basically mess with the source code in my own module that I'm putting the custom classes in) or create my own copy for each and make it return a DataFrame in the transform() method?
I am spoiled by the way R libraries deal with this sort of stuff and I want to be able to explain and visualize what is happening and why.
*** UPDATE
@VivekKumar
num_pipeline = Pipeline([
("num_selector", VariableSelector(variable_type = "numeric")),
("num_imputer", Imputer(strategy = "median", copy = False))
])