242 lines
5.5 KiB
Markdown
242 lines
5.5 KiB
Markdown
# GPU 推理部署指南
|
||||
|
|
|
|||
|
|
## 配置参数
|
|||
|
|
|
|||
|
|
在 `.env` 文件中设置:
|
|||
|
|
|
|||
|
|
```env
|
|||
|
|
# 推理设备: cpu / cuda / dml / tensorrt
|
|||
|
|
OCR_DEVICE=cpu
|
|||
|
|
|
|||
|
|
# GPU设备编号(多GPU时指定,默认0)
|
|||
|
|
OCR_DEVICE_ID=0
|
|||
|
|
|
|||
|
|
# CPU线程数(仅 cpu 模式生效)
|
|||
|
|
OCR_INTRA_THREADS=4
|
|||
|
|
OCR_INTER_THREADS=2
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 方案一:CUDA(NVIDIA 显卡)
|
|||
|
|
|
|||
|
|
适用:有 NVIDIA 独立显卡(GTX/RTX/Quadro/Tesla 等)。
|
|||
|
|
|
|||
|
|
### 1. 确认显卡支持 CUDA
|
|||
|
|
|
|||
|
|
```powershell
|
|||
|
|
nvidia-smi
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
确保 CUDA Version >= 11.8。
|
|||
|
|
|
|||
|
|
### 2. 安装 CUDA Toolkit 和 cuDNN
|
|||
|
|
|
|||
|
|
**方法A(推荐):直接安装 onnxruntime-gpu**
|
|||
|
|
|
|||
|
|
onnxruntime-gpu 已内置必要的 CUDA/cuDNN 依赖(Windows 上通过 DirectML 或 CUDA EP):
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
pip uninstall onnxruntime
|
|||
|
|
pip install onnxruntime-gpu
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**方法B:手动安装 CUDA + cuDNN**
|
|||
|
|
|
|||
|
|
1. 下载 [CUDA Toolkit 11.8](https://developer.nvidia.com/cuda-11-8-0-download-archive)
|
|||
|
|
2. 下载 [cuDNN 8.x for CUDA 11.x](https://developer.nvidia.com/cudnn)
|
|||
|
|
3. 安装后将 cuDNN 的 bin/lib/include 复制到 CUDA 安装目录
|
|||
|
|
|
|||
|
|
### 3. 配置 .env
|
|||
|
|
|
|||
|
|
```env
|
|||
|
|
OCR_DEVICE=cuda
|
|||
|
|
OCR_DEVICE_ID=0
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 4. 验证
|
|||
|
|
|
|||
|
|
启动服务后查看日志:
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
推理设备: CUDA (device_id=0)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
或用 Python 测试:
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
import onnxruntime as ort
|
|||
|
|
print(ort.get_available_providers())
|
|||
|
|
# 应包含: ['CUDAExecutionProvider', 'CPUExecutionProvider']
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 方案二:DirectML(Windows 任意 GPU)
|
|||
|
|
|
|||
|
|
适用:Windows 系统,任意显卡(NVIDIA/AMD/Intel 核显),无需安装 CUDA。
|
|||
|
|
|
|||
|
|
### 1. 安装 onnxruntime-directml
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
pip uninstall onnxruntime
|
|||
|
|
pip install onnxruntime-directml
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2. 配置 .env
|
|||
|
|
|
|||
|
|
```env
|
|||
|
|
OCR_DEVICE=dml
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 3. 验证
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
import onnxruntime as ort
|
|||
|
|
print(ort.get_available_providers())
|
|||
|
|
# 应包含: ['DmlExecutionProvider', 'CPUExecutionProvider']
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
> **注意**:首次推理时 ONNX Runtime 会编译 DML 算子,可能耗时 10-30 秒,后续推理恢复正常速度。
|
|||
|
|
|
|||
|
|
### 4. 核显部署机(无独显)完整流程
|
|||
|
|
|
|||
|
|
开发机与部署机分离时,部署机常为只有核显的轻薄本/迷你主机。核显同样支持 DirectML 加速(Intel UHD / Iris Xe、AMD Radeon 核显均满足),无需任何 NVIDIA/AMD 驱动层面的额外配置。
|
|||
|
|
|
|||
|
|
**前置条件**
|
|||
|
|
|
|||
|
|
| 项目 | 要求 |
|
|||
|
|
|------|------|
|
|||
|
|
| 系统 | Windows 10 1903(Build 18362)及以上(RapidOCR 会检查,不满足则自动回退 CPU) |
|
|||
|
|
| 显卡 | 支持 DirectX 12 的核显(现代 Intel/AMD 核显均满足) |
|
|||
|
|
| Python | 3.10+ |
|
|||
|
|
|
|||
|
|
**全新环境安装**
|
|||
|
|
|
|||
|
|
```powershell
|
|||
|
|
# 1. 安装项目依赖
|
|||
|
|
pip install -r requirements.txt
|
|||
|
|
|
|||
|
|
# 2. 安装 DirectML 后端(替换掉 rapidocr 自带的 CPU 版 onnxruntime)
|
|||
|
|
pip uninstall onnxruntime
|
|||
|
|
pip install onnxruntime-directml
|
|||
|
|
|
|||
|
|
# 3. 验证后端可用
|
|||
|
|
python -c "import onnxruntime as ort; print(ort.get_available_providers())"
|
|||
|
|
# 应输出: ['DmlExecutionProvider', 'CPUExecutionProvider']
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**配置 .env**
|
|||
|
|
|
|||
|
|
```env
|
|||
|
|
OCR_DEVICE=dml
|
|||
|
|
# OCR_INTRA_THREADS / OCR_INTER_THREADS 仅 cpu 模式生效,dml 模式无需设置
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**启动并验证**
|
|||
|
|
|
|||
|
|
```powershell
|
|||
|
|
uvicorn main:app --host 0.0.0.0 --port 8000
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
- 启动日志应出现 `推理设备: DirectML`
|
|||
|
|
- 首次识别会编译着色器,等待 30-60 秒属正常现象,之后恢复正常速度
|
|||
|
|
|
|||
|
|
**核显加速预期**
|
|||
|
|
|
|||
|
|
| 场景 | 说明 |
|
|||
|
|
|------|------|
|
|||
|
|
| 加速比 | 约为 CPU 的 1.5-3 倍(Intel Iris Xe 单张 A4 约 1-2s) |
|
|||
|
|
| CPU 较强的机器 | 核显与 CPU 共享内存带宽,加速比更接近下限 1.5x |
|
|||
|
|
| 双显卡机器 | DML 使用系统默认适配器(通常为独显);当前代码未将 `OCR_DEVICE_ID` 透传给 DML,该参数对 dml 模式不生效 |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 方案三:TensorRT(NVIDIA 显卡,极致性能)
|
|||
|
|
|
|||
|
|
适用:追求最快推理速度的 NVIDIA GPU 用户。
|
|||
|
|
|
|||
|
|
### 1. 安装依赖
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
pip uninstall onnxruntime
|
|||
|
|
pip install onnxruntime-gpu
|
|||
|
|
|
|||
|
|
# 安装 TensorRT
|
|||
|
|
pip install tensorrt
|
|||
|
|
# 或从 NVIDIA 官网下载: https://developer.nvidia.com/tensorrt
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2. 配置 .env
|
|||
|
|
|
|||
|
|
```env
|
|||
|
|
OCR_DEVICE=tensorrt
|
|||
|
|
OCR_DEVICE_ID=0
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 3. 首次运行
|
|||
|
|
|
|||
|
|
TensorRT 首次加载模型时需要构建 engine(耗时 30 秒到几分钟),之后会缓存到默认模型目录。
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 方案四:OpenVINO(Intel CPU/GPU)
|
|||
|
|
|
|||
|
|
如果你使用 Intel CPU(特别是第 10 代及更新),OpenVINO 比默认 ONNX Runtime CPU 更快。
|
|||
|
|
|
|||
|
|
### 1. 安装
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
pip install openvino
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2. 修改 `engine/ocr_engine.py`
|
|||
|
|
|
|||
|
|
在 `_build_device_params()` 中添加 OpenVINO 分支(目前代码模板已预留结构)。
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 性能预期
|
|||
|
|
|
|||
|
|
| 设备 | 单张 A4 耗时 | vs CPU | 适用场景 |
|
|||
|
|
|------|-------------|--------|---------|
|
|||
|
|
| CPU (i7-12700) | ~2-5s | 基准 | 低并发、开发调试 |
|
|||
|
|
| CUDA (RTX 3060) | ~0.3-0.8s | 3-8x | 生产环境 NVIDIA GPU |
|
|||
|
|
| DirectML (RTX 3060) | ~0.5-1.0s | 2-5x | Windows 任意 GPU |
|
|||
|
|
| DirectML (Intel Iris Xe) | ~1-2s | 1.5-3x | 轻薄本核显 |
|
|||
|
|
| TensorRT (RTX 3060) | ~0.15-0.4s | 5-15x | 高吞吐生产环境 |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 故障排查
|
|||
|
|
|
|||
|
|
### CUDA 模式下报错 "CUDAExecutionProvider not found"
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# 确认安装了 onnxruntime-gpu
|
|||
|
|
pip show onnxruntime-gpu
|
|||
|
|
|
|||
|
|
# 确认 CUDA 可用
|
|||
|
|
python -c "import onnxruntime; print(onnxruntime.get_available_providers())"
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### DirectML 下首次推理卡住
|
|||
|
|
|
|||
|
|
首次推理需要编译着色器,等待 30-60 秒即可。后续推理恢复正常。
|
|||
|
|
|
|||
|
|
### GPU 显存不足
|
|||
|
|
|
|||
|
|
服务使用整图 OCR。如果处理大图时出现 OOM:
|
|||
|
|
- 在调用服务前降低输入图片分辨率
|
|||
|
|
- 降低并发请求数量
|
|||
|
|
- 或降低 `.env` 中 `OCR_REC_BATCH_NUM` 的值
|
|||
|
|
|
|||
|
|
### 多 GPU 环境
|
|||
|
|
|
|||
|
|
设置 `OCR_DEVICE_ID` 指定使用哪张卡:
|
|||
|
|
|
|||
|
|
```env
|
|||
|
|
OCR_DEVICE=cuda
|
|||
|
|
OCR_DEVICE_ID=1 # 使用第2张GPU
|
|||
|
|
```
|