Training data of variable video length in sign language

Viewed 66

I am trying to extract keypoints from training data of variable video length in sign language using MediaPipe and use them for LSTM training.

Because the length of video data is variable ex) 2 sec, 7 sec, 3 sec...

I've been thinking of three methodologies, and I'm wondering which of these is the best or if there are other optimal ones.

(The video rate is 30 frames per second.)

  1. Extract only 30 frames from every video.
    After checking the frame length of the video before extraction.
    Divide the frame by 30 and extract a frame from the video for each unit. ex) 2sec => 2 frame, 4 frame.... 60 frame..

  2. Specify the maximum length and fill the empty space with numpy.zeros(keypoints.shape).
    ex) maximum_length = 8sec , random_video_length = 2sec // (8 - 2) * 30 Fill the frame with numpy.zeros(keypoints.shape).

  3. Specify the maximum length and fill the missing image by copying the keypoints extracted from the last frame.
    ex) maximum_length = 8sec , random_video_length = 2sec
    => Copy and fill in the last frame of the random video.

0 Answers
Related