CausalFlow.ai
Platform Performance
Petit AveLang
Blog Discuss your deployment ↗
Platform Performance

Developers

Petit AveLang Blog Discuss your deployment ↗
November 2025 Compiler Optimization

VeriLocc: Using LLMs to Allocate GPU Registers

VeriLocc uses an LLM to assign GPU registers across NVIDIA and AMD hardware, then verifies each result.

Read more
August 2025 GPU Optimization

Optimizing FP4 Mixed-Precision Inference on AMD GPUs

Learn how we developed Petit, a collection of optimized FP16/BF16 x FP4 mixed-precision GPU kernels for AMD GPUs, achieving 1.74x faster inference and up to 3.7x performance improvements.

Read more
CausalFlow.ai

High-performance inference for AMD GPUs.

© 2024–2026 CausalFlow Inc.