Basically, I want to transform a point from the image of camera 2 (x_pixel_c2, y_pixel_c2) to a pixel point (x_pixel_c2, y_pixel_c3) in camera 3.
For a simple setup (no distortion parameters, no rectification), I would usually:
- assign a certain distance to the point in C2
- compute the coordinates in 3D space of C2 from the intrinsics
- compute the World coordinates from the extrinsic matrix of C2
- compute the 3D coordinates in C3 from its extrinsic matrix
- project in pixel space of C2
For the Kitty dataset, I do not use this approach because of the distortion parameters (especially). In Kitty, we have the projection matrices for the 4 rectified cameras, which from what I understand, are the transformation matrices relating the 3D coordinates in Camera 0 to Camera X. However, from the tests I've made, this does not work:
pt_3d_cam0 = np.dot(inverse(P_rect_02), pt_cam_2)
pt_cam_3 = np.dot(P_rect_03, pt_3d_cam0)
I'm not sure the projection matrices move the referential back to camera 0 as it is mentioned. For instance, the points are closer to the expected coordinates if I add a translation in X equal to the amount of the baseline.
I'm not sure if I'm missing something. If anyone has encountered a similar problem with Kitty, it would be appreciated if you could help.
Thank you!