How to optimize your tensorflow model by using TensorRT?

Viewed 1068

These are the instruction to solve the assignments?

  1. Convert your TensorFlow model to UFF
  2. Use TensorRT’s C++ API to parse your model to convert it to a CUDA engine.
  3. TensorRT engine would automatically optimize your model and perform steps like fusing layers, converting the weights to FP16 (or INT8 if you prefer) and optimize to run on Tensor Cores, and so on.

Can anyone tell me how to proceed with this assignment because I don't have GPU in my laptop and is it possible to do this in google colab or AWS free account. And what are the things or packages I have to install for running TensorRT in my laptop or google colab?

2 Answers

so I haven't used .uff but I used .onnx but from what I've seen the process is similar.

According to the documentation, with TensorFlow you can do something like:

from tensorflow.python.compiler.tensorrt import trt_convert as trt
converter = trt.TrtGraphConverter(
    input_graph_def=frozen_graph,
    nodes_blacklist=['logits', 'classes'])
frozen_graph = converter.convert()

In TensorFlow1.0, so they have it pretty straight forward, TrtGraphConverter has the option to serialized for FP16 like:

converter = trt.TrtGraphConverter(
    input_saved_model_dir=input_saved_model_dir,
    max_workspace_size_bytes=(11<32),
    precision_mode=”FP16”,
    maximum_cached_engines=100)

See the preciosion_mode part, once you have serialized you can load the networks easily on TensorRT, some good examples using cpp are here.

Unfortunately, you'll need a nvidia gpu with FP16 support, check this support matrix.

enter image description here

If I'm correct, Google Colab offered a Tesla K80 GPU which does not have FP16 support. I'm not sure about AWS but I'm certain the free tier does not have gpus.

Your cheapest option could be buying a Jetson Nano which is around ~90$, it's a very powerful board and I'm sure you'll use it in the future. Or you could rent some AWS gpu server, but that is a bit expensive and the setup progress is a pain.

Best of luck!

Export and convert your TensorFlow model into .onnx file.

Then, use this onnx-tensorrt tool to do the CUDA engine file conversion.

Related