Machine setup (GPU)
Floyo runs on NVIDIA H100 NVL GPUs. Each H100 NVL delivers 3.9 TB/s of HBM3 memory bandwidth - over 2x the bandwidth of RTX PRO 6000 GPUs - which helps workflows run faster, especially for higher-res generations, larger models, and longer videos
| H100 NVL | RTX PRO 6000 | RTX 5090 | What it means (AI/ML workflows) |
|---|---|---|---|---|
Memory Type & Size | 94GB HBM3 | 96GB GDDR7 | 32GB GDDR7 | HBM3 vs GDDR7: HBM3 is “on-package” stacked memory built for extreme speed. GDDR7 is fast memory chips on the board, but with less bandwidth than HBM setups. |
Memory bandwidth | 3.938 TB/s | 1.792 TB/s | 1.792 TB/s | Bandwidth = how fast the GPU can read/write VRAM. Diffusion/video workflows move a ton of data every step. When bandwidth is the limiter, higher bandwidth directly translates to faster steps and better throughput (especially big images, long video, and multi-model workflows). H100 NVL has ~2.2x higher memory bandwidth than RTX PRO 6000 and RTX 5090. |
FP8 Tensor | 3,341 TFLOPS | 2,015 TFLOPS | 1,676 TFLOPS | Think of this as “top-end AI math speed” when a workflow/runtime can use FP8. If a workflow is FP8-optimized (quantized), higher FP8 mean more images/frames per minute. H100 leads by a lot here. |
FP16 Tensor | 1,671 TFLOPS | 1,008 TFLOPS | 838 TFLOPS | Most open-source image/video generation inference runs in FP16/BF16 today. This often tracks real-world speed for diffusion steps. H100 is materially ahead, which is why we standardize on H100 NVL for predictable, fast workflow runs. |
TF32 Tensor | 835 TFLOPS | 504 TFLOPS | 210 TFLOPS | Mostly relevant to training/fine-tuning and some mixed-precision training paths - less important for typical “run a workflow and generate” inference. |
Tensor TFLOPS are peak theoretical values shown with “sparsity” where applicable; real-world speed varies by model, precision/quantization, and which part of the workflow is the bottleneck (often memory bandwidth for diffusion/video).