How to compute vertex positions via large, dense matrix-vector multiplication on GPU?

Viewed 64

My vertex positions V∈ℝ³ⁿ should be the result of multiplication of a large, dense matrix B∈ℝ³ⁿˣᵏ with a vector z∈ℝᵏ:

V = B z

In my case, k≈300, so as I understand this is far too big to store the relevant rows Bᵢ for the ith vertex as vertex attributes.

Currently, I'm computing this multiplication in the vertex shader by setting z as a uniform, by packing B into a square texture then using a texelFetch and a for-loop. Something like:

uniform int n; // number of vertices
uniform int k; // size of z
uniform int s; // size of texture square where B is packed
uniform float z[512];
uniform sampler2D tex;
in float id; // index of vertex
out vec3 v;

void main()
{
  v = vec3(0,0,0);
  for(int j = 0;j < k; j++)
  {
    int index = int(id)*k+j;
    int si = index % s;
    int sj = int((index - si)/s);
    v = v + texelFetch(tex,ivec2(si,sj),0).xyz*z[j];
  }

On my MacBook Pro M1, this works reasonably well for n≈50,000 and k≈100. Increasing either, I start to get dropped frames.

Is doing this in a vertex shader a good idea?

My computation is similar to blendshapes. How are those typically computed on the GPU?

Ideally, I'd like to stick to opengl. If not, is there an opencl or some other way to best achieve this?

1 Answers

I'm a little confused about your shader. z is never used, but there's access to an array q. I assume this is a typo?

Leaving aside shader type, depending on whether you're compute or memory bound, you may also be leaving some possibly substantial optimisation on the table:

  1. Your si/sj computations use integer division/division remainder. Try to avoid those expensive calculations in your hot loop. Instead, perform those calculations once before the loop, then increment si on every iteration, followed by simple test if you've hit the end of the row if (si >= s) { si = 0; ++sj; } I know they say avoid branching in shaders, but simple conditionals like this will either not branch at all, or at least have a good chance of being cheaper than integer division. (No point guessing: measure which is faster.)

  2. You say your matrix is dense. Is your input vector z/q, though? If a lot of elements there are 0, it may be wiser to skip the parts of the matrix that would be multiplied by 0 altogether. Either by explicitly testing for 0 in the input vector elements, or by passing the input vector in as a list of nonzero elements. (Arrays containing nonzero values and their indices, number of nonzero values.) Whether this is worth it depends on the extent to which your performance is limited by memory bandwidth; if it's more than half of elements, it's definitely worth trying, but given how expensive it is to read your huge matrix, it's likely worth it even for a small number of zeroes.

  3. Finally, you don't specify how you arrive at your giant matrix B. If you can decompose its calculation at all, that may be worth doing if it reduces the amount of data you need to read in your vertex shader.

Shader types

One of the main things that should influence your decision of whether to use an expensive vertex shader or a compute shader is whether you are re-running the computation with the same values at all. If these are all single-use calculations, you might be able to stick with a vertex shader. If your z/q vector and B matrix are constant over multiple render passes, you'll want to cache the result.

A compute shader will also give you more control over parallelism, and you don't have to mess around with packing matrices into textures: just use arrays.

What I mean by more control over parallelism: Instead of performing each matrix multiplication slice sequentially, you can compute each vector element in a work-group, with each work-item in the group computing a few slices, and then the group performs a parallel reduction in local/group memory to accumulate the final result. This is often more efficient for equally occupying all your GPU's shader units.

Choice of API/Platform

If you're targetting macOS and other Apple platforms, you might want to consider using Metal instead of OpenGL and OpenCL as the latter 2 are marked deprecated by Apple, and tools (profiling, debugging, …) support is non-existent compared to Metal. (Yes this is incredibly annoying.) On other platforms, OpenGL offers compute shaders as well, but Apple stopped implementing new OpenGL features before compute shaders, so the only way to implement them in macOS is via OpenCL or Metal. (Using Metal should let you determine whether your shader's performance is compute or memory bound, for example, and so lets you better guide your optimisation.)

Related