The inference platform for AMD GPUs

More tokens
per dollar on AMD.

Deploy fast, production-ready model endpoints on AMD GPUs. CausalFlow combines an optimized inference engine, fleet control, and model-specific performance work in one platform.

Performance on AMD

Measured inference speedups.

Higher is better

NVFP4 microbenchmark Up to 3.7×

versus hipBLASLt on AMD MI300X

End-to-end inference Up to 2.2×

Kimi 2.5 benchmark on AMD MI350X

Proven with AMD Trusted by AMD.

Inference optimization for production AMD GPU deployments.

AMD

From capacity to token factory

A product for every layer of AMD inference.

Five connected products span workload policy, model serving, cluster control, GPU kernels, and infrastructure.

View platform details
CausalFlow products spanning the full AMD inference stack

For neoclouds and enterprises

Build production inference on AMD.

CausalFlow works with neoclouds and enterprises running AMD GPU fleets to improve throughput, latency, and cost per token.

Discuss your deployment