Skip to content

Feature guide

Memory and throughput

Model customization

Input types

  • Multimodal inputs cover image, audio, and video models.
  • FP8 vision attention can accelerate Qwen3 vision encoders on supported NVIDIA and AMD GPUs.
  • Encoder-decoder models support sequence-to-sequence tasks.
  • Pooling runners provide embeddings, classifications, rewards, and scores.

Operations