onnxruntime
View on GitHubONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Cross-platform ML inference and training accelerator. Runs models from PyTorch, TensorFlow, and classical ML libraries via ONNX, applying graph optimizations and hardware acceleration on CPU, GPU, or NPU. Widely used runtime for deploying fast, low-cost model inference.
Use Cases
Accelerating deep learning inference across CPU, GPU, and NPUDeploying PyTorch and TensorFlow models cross-platform via ONNXServing classical ML models (scikit-learn, XGBoost, LightGBM) in productionSpeeding up transformer training on multi-node NVIDIA GPUsEdge and mobile model deployment with hardware accelerationQuantizing and graph-optimizing models for lower latency and costRunning LLM inference locally with optimized runtime backends
Built With
- Language
- C++
- Frameworks
- ONNX · PyTorch · TensorFlow · Keras · scikit-learn · LightGBM · XGBoost · CUDA · TensorRT · OpenVINO · DirectML · ROCm · CoreML · NNAPI · NumPy
Tags
onnx · inference-engine · model-serving · hardware-acceleration · cross-platform · quantization · graph-optimization · deep-learning · machine-learning · training · edge-deployment · cuda · tensorrt · cpp · python · performance