Use Cases
720p 30fps Fast Video Inference On Single Or Multi-gpu
1 project
Accelerate Llm Inference Throughput And Latency
1 project
Accelerating Deep Learning Inference Across Cpu, Gpu, And Npu
1 project
Accelerating Tts Inference With Vllm Or Tensorrt-llm
1 project
Add A Standalone KV Cache Daemon To Existing Inference Engines
1 project
Add Local Llm Inference To Existing Sdk-based Applications
1 project
Analyze Llm Inference Stack Performance With Aiperf
1 project
Batch Inference Over Software Tasks
1 project
Benchmark Disk-streaming Vs Resident Inference Performance
1 project
Benchmark Inference Latency And Throughput
1 project
Benchmark Inference Performance Across Optimized Models
1 project
Benchmark Inference Speed And Concurrency
1 project
Benchmarking Inference Latency And Throughput Vs Upstream Pytorch Ports
1 project
Benchmarking Inference Throughput And Latency
1 project
Build Lightweight C/c++ Inference Into Apps And Edge Devices
1 project
Build Python Model Pipelines For Inference
1 project
Cli Inference For Batch Audio Conversion
1 project
Cpu Inference Of Sharded Models
1 project
Cpu-only Inference With No Gpu
1 project
Cpu-only Local Llm Inference Without Gpu Or Blas
1 project
Custom-metrics Autoscaling For Inference Services
1 project
Cut Inference Cost For Coding Agents Without Changing Agent Code
1 project
Deploying A Tts Inference Server Via Grpc/fastapi/docker
1 project
Disk-backed Local Cluster Inference Across Machines
1 project
Distributed Inference And Fine-tuning Across Macs
1 project
Distributed Inference Over Rpc Backend
1 project
Distributed Multi-gpu/multi-node Inference
1 project
Distributed Multi-mac Inference Over Ring/thunderbolt
1 project
Distributed Multi-node Inference Cluster
1 project
Edge Inference
1 project
Export Mask Decoder To Onnx For In-browser Inference
1 project
Fast Low-latency Per-decision Inference (e.g. Game Ai)
1 project
Gpu-accelerated Local Inference With Nvidia Cuda
1 project
High-throughput Batch Image Generation With Vllm Offline Inference
1 project
High-throughput Llm Inference And Serving
1 project
Hybrid Cpu+gpu Inference For Models Larger Than Vram
1 project
Improve Moe Inference Throughput
1 project
Indicate Model Inference Or Loading
1 project
Inference On Apple Silicon Macs
1 project
Inference Routing To Self-hosted Models On Kubernetes
1 project
Integrate Diffusion Inference Into C/c++ Applications
1 project
Local Inference Without Sending Data To A Hosted Api
1 project
Local Offline Llm Inference With Ollama/vllm/lm Studio
1 project
Local Private Llm Inference Without Api Providers
1 project
Local-first Inference With Llama.cpp
1 project
Long-context Inference With Compressed Mla KV State
1 project
Low-latency Chat/completion Inference At Scale
1 project
Machine-learning Inference
1 project
Memory-efficient Inference On Long Documents Via Compact Kv-cache Merging
1 project
Multi-node Distributed Inference With Tensor/expert Parallelism
1 project
Multimodal Image/video/audio Chat Inference
1 project
Neural Vocoder Inference
1 project
Offline Inference With No Api Keys Or Accounts
1 project
Offline Inference Without Cloud Api Or Output Tokens
1 project
Offline On-device Inference Via Llama.cpp With Lora Adapters
1 project
Offline Or Air-gapped Inference
1 project
On-device Inference Via Llama.cpp-omni
1 project
Point Claude, Codex, And Opencode At Custom Inference Endpoints Or Byo Models
1 project
Prefill-decode Disaggregated And Multi-node Inference
1 project
Prefill/decode Disaggregation And Dp-aware Routing For Inference Engines
1 project
Private On-prem Multi-modal Inference
1 project
Quantize Models For Faster Inference And Smaller Checkpoints
1 project
Quantized Inference (fp4/fp8/int4/awq/gptq)
1 project
RL Rollout Infrastructure Spanning Agent, Inference And Training
1 project
Recursive Inference Over Inputs Larger Than The Model Context Window
1 project
Route Tasks To Best-fit Models Through An Optional Inference Router
1 project
Run 70b Llm Inference On A Single 4gb Gpu
1 project
Run A First Local Llm Inference With Ollama
1 project
Run AI Inference Close To End Users
1 project
Run Bedrock Batch Inference Jobs Via Step Functions
1 project
Run Inference Through Cli Tools Like Claude Or Gemini
1 project
Run Inference/training Through A Local Gradio Webui
1 project
Run Local Image Inference On Mac Via Sd.cpp
1 project
Run Local Llm Inference For Music-production Workflows
1 project
Run Local On-device Llm Inference
1 project
Run Local Video Inference Via Wan2gp Server
1 project
Run On-device Inference With Local Models
1 project
Run Private/offline Inference Without Sending Data To Cloud Apis
1 project
Run Quantized Llm Inference In The Browser Via Onnx + Wasm
1 project
Run Uncensored 27b Inference Locally On Apple Silicon
1 project
Running Fully Local/private Inference Via Privacy Mode Or Byok Providers
1 project
Running Llm Inference Locally With Optimized Runtime Backends
1 project
Self-host An Openai-compatible Inference Server
1 project
Serve An Openai-compatible Local Inference Server
1 project
Serve Batched State/question/candidate Requests Over An Http Inference Api
1 project
Speed Up Inference 3x With Block-wise Quantization
1 project
Start/stop The Local LM Studio Inference Api Server
1 project
Stream Model Inference Logs To The Terminal
1 project
Study A From-scratch Transformer/moe Inference Engine In Portable C99
1 project
Summarizing Very Large Logs With Budgeted Fan-out Inference
1 project
Sync And Async Inference Sessions Over Websocket
1 project
Unwrapping Lightning Checkpoints For Inference
1 project
Use The Gateway As A Unified Authenticated Proxy For Model Inference Endpoints
1 project