Import the output of an R script with Python

Viewed 592

my first post in this stackoverflow! :)

I am trying to import a table (output of an R script) using Python. This would be very helpful to avoid to translate a huge script designing a complex data.table in R.

Righ now I know how to call the R script with Python using the following code:

import os 
import subprocess
 
#Launch selected script
command = 'C:/Program Files/R/R-3.4.0/bin/x64/Rscript.exe'
path2script = 'C:/mypath/myscript.R'
cmd = [command, path2script]
a = subprocess.call(cmd)

But then I dont know how to use the table, output of the R code, using my Python script. Would you have any idea?

Many thanks

EDIT:

I tried the solution from @punter below

import subprocess

with subprocess.Popen(['/command/to/run', '/other/parameters'], stdout=subprocess.PIPE) as proc:
    table = proc.stdout.read()

But then the table as a strange format like this: (it is a subset)

A\r\n COL1 COL2 COL3\r\n 1: 2015-06-17 05:19 NA <NA>\r\n 2: 2015-06-17 05:19 NA <NA>\r\n 3: 2015-06-17 05:19 NA <NA>\r\n 4: 2015-06-17 05:19 NA <NA>\r\n 5: 2015-06-17 05:19:29 NA <NA>\r\n 

and when I try the code below I get all the content in the column names

s=str(table) 
data = StringIO(s) 
df=pd.read_csv(data)

[0 rows x 111 columns]

EDIT NUMBER 2

trying with this "ISO-8859-1" in str like str(table, "ISO-8859-1") seems to be working and I could notice that I had more than one table in my script. I am rerunning everything in a clean way I hope it will work! :)

4 Answers

For any process in general, the stdout can be captured as follows:

import subprocess

with subprocess.Popen(['/command/to/run', '/other/parameters'], stdout=subprocess.PIPE) as proc:
    table = proc.stdout.read()

this is a blocking call, this means that the execution will halt at table = proc.stdout.read().

Note The above code may not work as required for large tables, as the process will block at the line shown above for a very long time.

Consider adjusting R script to use write.csv() as last line without a file name to dump data to console. This will allow string values to be quoted. Then in Python, use subprocess.Popen to receive output with byte handling for migration into Pandas data frame.

R

...

write.csv(my_r_df, file="", row.names=FALSE)

Python

import subprocess
from io import StringIO
import pandas as pd
 
# RUN R SCRIPT
command = r'C:\Path\To\Rscript.exe'
path2script = r'C:\Path\To\R\Code.R'

a = subprocess.Popen([command, path2script], 
                      stdin=subprocess.PIPE, 
                      stdout=subprocess.PIPE, 
                      stderr=subprocess.PIPE)
                      
output, error = a.communicate()

# IMPORT PANDAS DATA FRAME
my_pandas_df = pd.read_csv(StringIO(output.decode('utf-8')))

my_pandas_df

Of course, too, you can have R simply write data to .csv and then import into Pandas.

You can copy the R dataframe to the clipboard and then load it to Pandas using read from clipboard. Here is how it works.

Use this code to copy a R dataframe to clipboard, then go to Jupyter and use pd.read_clipboard().

# If you are on Windows, use:
# R-Studio
write.table(df, "clipboard", sep="\t", col.names=TRUE)

# Jupyter
df = pd.read_clipboard(sep="\t")

# On Mac, use:
# R-Studio
clip <- pipe("pbcopy", "w")                       
write.table(df, file=clip)                               
close(clip)

# Jupyter Notebook
df = pd.read_clipboard()
Related