Converting a list of R,B,G arrays into color names

Viewed 372

I am trying to convert RBG values from an image into color names. The function I am using outputs the RBG values as the following:

get_colors(image, 3, False) → [array([180.60849865, 187.12872355, 191.33901941]), array([254.80780926, 254.83214586, 254.8409425 ]), array([191.33977471, 197.17187803, 200.97835918])]

I would like to iterate these variables as individual arguments in pairs of three through the following function:

def recognize_color(R,G,B):
    minimum = 10000
    for i in range(len(csv)):
        d = abs(R- int(csv.loc[i,"R"])) + abs(G- int(csv.loc[i,"G"]))+ abs(B- int(csv.loc[i,"B"]))
        if(d<=minimum):
            minimum = d
            cname = csv.loc[i,"color_name"]
    return cname

I am mostly struggling with this list of arrays and permutating them into the function. Ideally I want to create a list of color names from the list of RGB arrays

2 Answers

You can use a comprehension with argument unpacking:

[recognize_color(*color) for color in get_colors(image, 3, False)]

At each iteration, color will be an array with three elements, which can be expanded as any iterable into a tuple of arguments to recognize_color.

The result will be a list of names, one for each return value of get_colors.

You can also use itertools.starmap for this:

result = itertools.starmap(recognize_color, get_colors(image, 3, False))

This results in a lazy generator that you can iterate over to get the values on demand, including turning it into a list:

list(result)

An Alternative to for loops: Vectorization

This answer is a bit different. Typically, if you can vectorize your operations, the execution time could drop significantly.

Here is a solution which processes the whole array of [R, G, B] points, without using any for loop.

Steps to reproduce:

  • Run section C
  • Run section B
  • Run section A
  • Finally run the topmost code block.

Or, you could run this notebook:

  • Open in Colab
# This will return a pandas.Series of the 
# color-names for each color
recognize_colors(dfColors = dfc, 
                 dfColorRef=dfr, 
                 target_column="COLOR", 
                 color_cols=["R", "G", "B"], 
                 threshold=10_000)

A. Vectorized Code: recognize_colors()

You will only need this function to get what you want, without using any for loop.

from pandas import DataFrame
from typing import List

def recognize_colors(dfColors: DataFrame, 
                     dfColorRef: DataFrame, 
                     target_column: str="COLOR", 
                     color_cols: List[str] = ["R", "G", "B"],
                     threshold: int= 10_000):
    from numpy import abs
    ### The following four lines explain what is happening here 
    #   Here we use numpy array broadcasting to calculate ``d``
    #   in the foloowing four steps.
    # dR = np.abs(dfc.R.values.reshape(-1, 1) - df.R.values.reshape(1, -1))
    # dG = np.abs(dfc.G.values.reshape(-1, 1) - df.G.values.reshape(1, -1))
    # dB = np.abs(dfc.B.values.reshape(-1, 1) - df.B.values.reshape(1, -1))
    # d = dR + dG + dB
    #----------------------------------------------------------------------
    ## This is a condensed one-line construction of ``d`` using 
    #  numpy array broadcasting.
    d = abs(dfColors[color_cols].values.reshape(-1,1,3) - 
               dfColorRef[color_cols].values.reshape(1,-1,3)).sum(axis=2)
    return dfColorRef.loc[(d * (d <= threshold)).argmin(axis=1), target_column]

B. Dummy Data

We will use the following 5 rows of color as our dummy data to test the solution.

|   R |   G |   B |
|----:|----:|----:|
|  40 | 149 | 225 |
|  20 | 140 | 250 |
| 141 | 205 | 135 |
| 165 |  25 | 136 |
|  64 |  49 |  81 |

Now let's create the dummy data and load it in two dataframes:

  • dfc: dataframe of colors you want names for
  • dfr: dataframe of colors with names in the target_column, COLOR.
import numpy as np
import pandas as pd

input_colors = np.array([
    [ 40, 149, 225],
    [ 20, 140, 250],
    [141, 205, 135],
    [165,  25, 136],
    [ 64,  49,  81]
])
# Create a dataframe with all the colors
dfc = pd.DataFrame(input_colors, columns=["R", "G", "B"])

# Create another dataframe with some reference columns
#       While creating this I assume column "COLOR" is 
#       the name if the color.
# Note: In your case, you will read-in this information 
#       from some ".csv" file.
dfr = reference_colors(colors=None, show_progressbar=False)

C. Dummy Data Code

All the code below is only necessary to create dummy data and not required for the main solution.

#@title **Dummy Data Functions: `RUN`**
# import numpy as np
# import pandas as pd
# from tqdm.notebook import tqdm_notebook as tqdm
# from typing import Iterable, Optional, Union, List

def prepare_random_colors(
    seed: int=1234, 
    step: int=10, 
    offset_delta: int=0, 
    use_rgba: bool=False, 
    alpha: Optional[Union[int, float]] = None) -> np.ndarray:
    """Returns an array of colors of shape: (N, 3) or (N, 4).
    
    Logic:
        - An array ``base`` is created with values between 
          [0, 255] at steps of ``step``.
        - Another array ``offset`` of shape as shown below 
          is created randomly:
          - Shape:
            - (N, 3): for rgb colors
            - (N, 4): for rgba colors
          - N = base.shape[0]
          - The highest value

    Params:
        seed (int): Random number generator seed. 
                    - Constraint: ``seed >= 0``
                    - Default: `1234`
        step (int): Steps at which base-colors are generated. 
                    - Constraint: ``step >=1``
                    - Default: `10`
        offset_delta (int): An integer (+ or -). 
                            - Recommended: ``abs(offset_delta) <= step``
                            - Default: `0`
        use_rgba (bool): A flag on whether to use "rgba". When ``False``, 
                         "rgb" colors are generated.
                     - Default: `False`
        alpha (int, float, None): If you want to use a fixed alpha value in 
                                  rgba colors (must set ``use_rgba`` to ``True``),
                                  set ``alpha`` to a fixed value (int or float): 
                                  - ``[0., 1.]`` (int) 
                                  - ``[0, 255]`` (float)
                                  If, ``alpha`` = ``None``, and if ``use_rgba`` = ``True``,
                                  then randomly generated alpha values are used.

    Example:

        colors = prepare_colors(seed=1234)

    """
    rng = np.random.RandomState(seed)
    base = np.arange(0, 255, step)
    cols = 4 if use_rgba else 3
    offsets = rng.randint(0, step + offset_delta, size=(base.size, cols))
    colors = np.clip(base.reshape(-1, 1) + offsets, 0, 255)
    if (alpha is not None) and use_rgba:
        colors[:,:,:,-1] = alpha
    return colors

def clamp(x): 
  return max(0, min(x, 255))

def rgb2hex(r, g, b):
    r, g, b = clamp(r), clamp(g), clamp(b)
    return "#{0:02x}{1:02x}{2:02x}".format(r, g, b)

def rgba2hex(r, g, b, a):
    r, g, b = clamp(r), clamp(g), clamp(b), clamp(a)
    return '#{:02x}{:02x}{:02x}{:02x}'.format(r, g , b, a)

def reference_colors(colors: Optional[np.ndarray]=None, 
                     demoseed: int=1234, 
                     show_progressbar: bool=False):
    """Creates a reference-colors' dataframe from a numpy array of colors.

    Params:
        colors (numpy.ndarray): Array of colors. Shape: (num_rows, 3)
        demoseed (int): A seed for generating demo data for colors. 
                        This is used when ``colors = None``.

    Example:

        # colors = prepare_colors(seed=1234)
        df = reference_colors(colors=None, show_progressbar=False)

    """
    if colors is None:
        colors = prepare_random_colors(seed=demoseed)
    df = pd.DataFrame(colors, columns=["R", "G", "B"])
    if show_progressbar:
        # source: https://towardsdatascience.com/progress-bars-in-python-and-pandas-f81954d33bae 
        tqdm.pandas()
        df["HEX"] = df.progress_apply(lambda row: rgb2hex(*row), axis=1).str.upper()
    else:
        df["HEX"] = df.apply(lambda row: rgb2hex(*row), axis=1).str.upper()
    df["COLOR"] = df["HEX"]
    return df
Related