Currently, I am trying to create a pydantic model for a pandas dataframe. I would like to check if a column is unique by the following
import pandas as pd
from typing import List
from pydantic import BaseModel
class CustomerRecord(BaseModel):
id: int
name: str
address: str
class CustomerRecordDF(BaseModel):
__root__: List[CustomerRecord]
df = pd.DataFrame({'id':[1,2,3],
'name':['Bob','Joe','Justin'],
'address': ['123 Fake St', '125 Fake St', '123 Fake St']})
df_dict = df.to_dict(orient='records')
CustomerRecordDF.parse_obj(df_dict)
I would now like to run a validation here and have it fail since address is not unique.
The following returns what I need
from pydantic import root_validator
class CustomerRecordDF(BaseModel):
__root__: List[CustomerRecord]
@root_validator(pre=True)
def unique_values(cls, values):
root_values = values.get('__root__')
value_set = set()
for value in root_values:
print(value['address'])
if value['address'] in value_set:
raise ValueError('Duplicate Address')
else:
value_set.add(value['address'])
return values
CustomerRecordDF.parse_obj(df_dict)
>>> ValidationError: 1 validation error for CustomerRecordDF
__root__
Duplicate Address (type=value_error)
but i want to be able to reuse this validator for other other dataframes I create and to also pass in this unique check on multiple columns. Not just address.
Ideally something like the following
from pydantic import root_validator
class CustomerRecordDF(BaseModel):
__root__: List[CustomerRecord]
_validate_unique_name = root_unique_validator('name')
_validate_unique_address = root_unique_validator('address')