I have this dataframe (df)
|id | binary_col |
+-------------------------------+
|1 | [00 00 00 00 00 00 00 01] |
|2 | [00 00 00 00 00 00 00 01] |
|3 | [00 00 00 00 00 00 00 01] |
|4 | [00 00 00 00 00 00 00 02] |
|5 | [00 00 00 00 00 00 00 02] |
with these schema (df.printSchema())
|-- id: int (nullable = true)
|-- binary_col: binary (nullable = true)
And I want to filter only the values with [00 00 00 00 00 00 00 02]
I've tried:
from pyspark.sql import functions as F
- df.filter(F.col('binary_col')=='[00 00 00 00 00 00 00 02]')
- df.filter(F.col('binary_col')=='00 00 00 00 00 00 00 02')
but none of them worked.
When I try
df.filter(F.col('binary_col')==True) I get the error that cannot resolve '(`binary_col` = true)' due to data type mismatch: differing types in '(`binary_col` = true)' (binary and boolean).;
Any thoughts?