produce an equation for this { 20702 20709 20695 20703} to produce this {20714}

Viewed 70

I'm new here and actually this is my first question so please bear with me if it's not a great question.

Can anyone produce an equation to predict results column using any other data from the above lines?

if it could be done though i appreciate the help

i was given a xlsl file with 5 columns containing these numbers here is a slice of it

A.      B        C       D.     Result
20689   20724   20689   20702   20703
20702   20709   20695   20703   20714
20703   20714   20700   20714   20714
20714   20717   20702   20714   20728
20713   20732   20709   20728   20717
20728   20734   20714   20717   20692
20717   20734   20688   20692   20712
20692   20723   20683   20712   20705
20713   20714   20670   20705   20714
20704   20721   20692   20714   20714
20714   20714   20692   20714   20712
20714   20723   20707   20712   20726
20712   20726   20701   20726   20724
20726   20733   20720   20724   20724
20724   20724   20722   20724   20735
20724   20740   20722   20735   20736
20735   20738   20730   20736   20686
20736   20736   20682   20686   20722
20686   20728   20682   20722   20727
20722   20732   20720   20727   20705
20727   20732   20702   20705   20705
20705   20717   20702   20705   20709
20705   20715   20702   20709   20721
20709   20721   20700   20721   20718
20721   20731   20716   20718   20711
20716   20717   20711   20711   20691
20712   20713   20690   20691   20690
20691   20695   20687   20690   20717
20690   20717   20690   20717   20727
20717   20732   20712   20727   20727
20726   20733   20719   20727   20708
20727   20727   20707   20708   20692
20708   20710   20686   20692   20673
20692   20694   20673   20673   20681
20671   20693   20667   20681   20691
20681   20691   20675   20691   20666
20689   20689   20662   20666   20689
20666   20695   20666   20689   20708
20689   20723   20689   20708   20688
20708   20708   20686   20688   20677
20688   20689   20672   20677   20666
20677   20678   20662   20666   20681
20666   20681   20655   20681   20668
20685   20685   20663   20668   20647
20672   20672   20647   20647   20656
20647   20675   20643   20656   20638
20656   20665   20638   20638   20646
20638   20646   20628   20646   20623
20646   20661   20608   20623   20642

any help is immensely appreciated

3 Answers

Its not the solution to your problem, but maybe, this code snippet helps you nonetheless:

import numpy as np
import pandas as pd
df = pd.read_excel("data/data.xlsx")
df["Result"] = (df["A"] + df["B"] + df["C"] + df["D"]) // 4
df

you could think about it and try to find a pattern how the result depend on the columns (A,B,C,D)

how ever if this didn't success , there is a lot of solutions to these kind of problems in machine learning, so one could assume a linear relation between the features ( columns A,B,C,D) and the results , and try to find the parameter of an equation

Ax + By + Cz + Dw + M = result

and there's a lot more algorithm to use, if you knew about machine learning field .

A typical method is to assume that the value in Result can be obtained from the four values in A, B, C, D through some formula that involve parameters; then use an algorithm to find the parameters that best fit your data. The formula you chose is usually called a model.

This process is called regression.

In the particular case where the model looks like a weighted sum of the parameters, then this is called linear regression. This is the easiest case and the algorithms to find the best parameters are quite simple. I insist here on the fact that the term "linear" in linear regression concerns the parameters of the model, not the values from the data. For instance, if you assume that Result can be written as a polynomial in the four variables A, B, C, D, then you can use linear regression to find the coefficients of the polynomial - the model is linear in the coefficients of the polynomial, even though it is not linear in the variables A, B, C, D.

Note that the overall success of the method depends entirely on the model you chose to begin with. There are long discussions to be made about the choice between a simple model and a complicated model, and regularization techniques to allow a complicated model if necessary, but tell the optimization algorithm that if possible you would prefer a simpler model.

I won't go into more details about regularization techniques; here is a simple code using the python module sklearn.linear_model.LinearRegression to express Result as an affine combination of A, B, C, D.

import sklearn.linear_model as sklin
import pandas as pd

data = pd.read_csv('data.csv')

lm = sklin.LinearRegression()
lm.fit(data[['A', 'B', 'C', 'D']], data['Result'])

print(lm.coef_)
# array([ 0.1145072 ,  0.47290074, -0.36769957,  0.74233087])
print(lm.intercept_)
# 774.4947813684666

data['Predicted'] = lm.predict(data[['A', 'B', 'C', 'D']])
print(data)
#         A      B      C      D  Result     Predicted
# 0   20689  20724  20689  20702   20703  20704.326408
# 1   20702  20709  20695  20703   20714  20697.257624
# 2   20703  20714  20700  20714   20714  20706.063777
# 3   20714  20717  20702  20714   20728  20708.006659
# 4   20713  20732  20709  20728   20717  20722.804398
# ...

What did I do? I asked sklearn.linear_model.LinearRegression to find the best parameters a, b, c, d, e such that Result ≈ a * A + b * B + c * C + d * D + e; and the answer was [a, b, c, d] = [ 0.1145, 0.4729, -0.3677, 0.7423] and e = 774.49. I added the predictions as an extra column in the dataframe; the relation is:

Predicted = 0.1145 * A + 0.4729 * B - 0.3677 * C + 0.7423 * D + 774.49

Can we do better? You can try with a more complicated model, and see if you get better predictions.

Related