I created binary_cross_entropy_loss according to formula below:
import numpy as np
def binary_cross_entropy_loss(y_hat: np.ndarray, y_true: np.ndarray) -> float:
y_hat = np.clip(y_hat, 1e-7, 1 - 1e-7)
return -(y_true * np.log(y_hat) + (1 -y_true) * np.log(1 - y_hat)).mean()
However when I tried to compare TensorFlow BinaryCrossentropy with from_logits=False I saw that results are different but with from_logits=True are nearly identical.
My questions are:
- Where does this difference come from?
- How to achieve quite close results with from_logits=False?
- Which formula with from_logits=False is better and why?
Code below:
import tensorflow as tf
y_true: np.ndarray = np.float32([0, 1, 0, 0])
y_pred: np.ndarray = np.float32([-18.6, 0.51, 2.94, -12.8])
print("numpy, from_logits=False: ", binary_cross_entropy_loss(y_pred, y_true))
print("tensorflow, from_logits=False: ", tf.keras.losses.BinaryCrossentropy(from_logits=False)(y_true, y_pred).numpy())
def sigmoid(x: np.array) -> np.ndarray:
return 1 / (1 + np.exp(-x))
print("numpy, from_logits=True: ", binary_cross_entropy_loss(sigmoid(y_pred), y_true))
print("tensorflow, from_logits=True: ", tf.keras.losses.BinaryCrossentropy(from_logits=True)(y_true, y_pred).numpy())
numpy, from_logits=False: 4.1539326
tensorflow, from_logits=False: 4.0016456
numpy, from_logits=True: 0.86545783
tensorflow, from_logits=True: 0.865458