Modern C++ and Python Script to Accelerate and Optimize Artificial Intelligence Models and Deep Neural Networks via TensorRT on CUDA device, and Deploy the Models for Image Processing, Computer Vision and Machine Vision, such image Super-Resolution, YOLO serials detection.
- [x]YOLO26 ---> TensorRT Deploy GPU Accelerate inference
- [x]YOLO v11 model ---> TensorRT Deploy GPU Accelerate inference
- [x]Pytorch model ---> ONNX graph ---> TensorRT Engine ---> Deploy GPU Accelerate
- [x]TensorRT
- [x]CUDA kernel function
- [x]Convert Pytorch and TensorFlow models into Onnx
- [x]C++ implementation of YOLOv11 model using TensorRT API from ultralytics
git clone DeepLearningDeployment
cd DeepLearningDeployment
cmake -S . -B build -G "Visual Studio 17 2022" -DCMAKE_BUILD_TYPE:STRING=Debug
cmake -S . -B build -DCMAKE_BUILD_TYPE:STRING=Debug
cmake --build build --config Debug
cmake -S . -B build -G "Visual Studio 17 2022" -DCMAKE_BUILD_TYPE:STRING=Release
cmake -S . -B build -DCMAKE_BUILD_TYPE:STRING=Release
cmake --build build --config Release
cmake --build build --target clean
cmake --install build --prefix ./install
# 自动以最大线程数进行并行编译
sudo cmake --build build --target all -j12
# YOLOv11
pip install --upgrade ultralytics
. DeepLearningDeployment
|—— MNIST
| |—— README.md
|—— SISR
| |—— README.md
|—— TensorRT
| |—— CUDA_DriverAPI
|—— | |—— CMakeLists.txt
| |—— CUDA_RuntimeAPI
|—— | |—— CMakeLists.txt
| |—— TensorRT_Basic
|—— | |—— CMakeLists.txt
| |—— README.md
|—— external
| |—— TensorRT
| | |—— bin
| | |—— python
| | |—— lib
| | |—— include
| |—— LogModule
| | |—— lib
| | |—— include
| |—— Utility
| | |—— lib
| | |—— include
| |—— README.md
|—— Application_Python
| | |—— README.md
| | |—— requirements.txt
|—— Application_Cpp
| | |—— README.md
| | |—— CMakeLists.txt
|—— CMakeLists.txt
|—— requirements.txt
|—— .clang-format
|—— .gitignore
|—— README.md
- CUDA Driver and Runtime API
- 利用TensorRT加速深度神经网络模型
- Using TensorRT to Accelerate Deep Neural Network Models


