World co-ordinate point cloud generation from DEPTH16

Viewed 213

On an android device with a separate camera sensor and depth sensor, I'm trying to map the DEPTH16 output to a world co-ordinate point cloud usable in ARcore. The main calculation needed to do this is to map the depth output to the coordinate system of the primary camera (and ARcore can translate from there).

As I understand it, I need to apply the depth camera intrinsics (focal length, principal point offsets, etc), then apply the lens pose rotation to the primary camera coodinate system, then apply the translation which accounts for the fact that the depth sensor sits a few mm away from the camera lens. This is mostly documented in the Camera Characteristics API documentation, but doesn't seem to work for me in practice.

Here's what I'm doing:

  1. Pull the intrinsics, lens pose rotation and lens pose translaton from CameraCharacteristics of the depth sensor. This allows to populate the following parameters:
    public class Intrinsics {
        float width;
        float height;
        float focalLength;
        // intrinsic matrix:
        float f_x;
        float f_y;
        float c_x;
        float c_y;
        float s;
        // lens pose rotation
        float r_x;
        float r_y;
        float r_z;
        float r_w;
        // lens pose translation
        float t_x;
        float t_y;
        float t_z;
    }
  1. Apply the instrinsic transformation. From the sensor co-ordinate system, let's say my bottom-right pixel registers a depth of 1m. This gives it co-ordinates of [239, 179, 1.00] for x, y, z.

If I do the following arithmetic based on the focal lengths and principal points, I seem to get outputs that seem correct:

double[] intrinsicAdjusted = new double[]{
  (x - i.c_x) * depth / i.f_x,
  (y - i.c_y) * depth / i.f_y,
  depth
};

However, I've also tried using the transformation matrix suggested in the Camera2 docs (but the inverse as I want camera to world, not world to camera):

Mat point = new Mat(3, 1, CvType.CV_32F);
point.put(0, 0, (double)x, (double)y, depth);
Mat K = new Mat(3, 3, CvType.CV_32F);
K.put(0,0,
  i.f_x, i.s,   i.c_x,
  0,     i.f_y, i.c_y,
  0,     0,     1.0);
Mat Kinverse = K.inv();
Mat instrinsicAdjusted = new Mat(1, 3, CvType.CV_32F);
Core.gemm(Kinverse, point, 1, new Mat(), 0, instrinsicAdjusted, 0);

The skew is zero, so I think the approaches should work out the same, but they do not.

  1. Apply the rotation suggested by API docs to change the co-ordinates to that of the primary camera. I get very bogus results from the following (which uses OpenCV):
Mat point = new Mat(3, 1, CvType.CV_32F);
point.put(0, 0, instrinsicAdjusted[0], instrinsicAdjusted[1], instrinsicAdjusted[2]);
Mat R = new Mat(3, 3, CvType.CV_32F);
R.put(0, 0,
  1 - 2*(i.r_y*i.r_y) - 2*(i.r_z*i.r_z), 2*i.r_x*i.r_y - 2*i.r_z*i.r_w, 2*i.r_x*i.r_z + 2*i.r_y*i.r_w,
  2*i.r_x*i.r_y + 2*i.r_z*i.r_w, 1 - 2*i.r_x*i.r_x - 2*i.r_z*i.r_z, 2*i.r_y*i.r_z - 2*i.r_x*i.r_w,
  2*i.r_x*i.r_z - 2*i.r_y*i.r_w, 2*i.r_y*i.r_z + 2*i.r_x*i.r_w, 1 - 2*i.r_x*i.r_x - 2*i.r_y*i.r_y);
Mat rotated = new Mat(1, 3, CvType.CV_32F);
Core.gemm(R, point, 1, new Mat(), 0, rotated, 0);
  1. If things didn't go awry in the rotation, presumably the translation would work:
Mat translated = new Mat(3, 1, CvType.CV_32F);
translated.put(0, 0, i.t_x, i.t_y, i.t_z);
Core.add(translated, rotated, translated);

Thus, is this overall approach correct? Why do I get two different results for the instrinsic adjustment? And why does the rotation not seem to work?

0 Answers
Related