Porting from TritonAll pages
Guides
Porting from Triton
Most Triton kernels port by changing the imports and the device. This page lists the steps, what differs, and the Triton features that Enceladus refuses.
Port a kernel
To port a kernel, complete the following steps:
- Replace
import triton.language as tlwithimport enceladus.language as tl, and@triton.jitwith@enceladus.jit. - Move tensors from
cudatomps. - Mask the M and N edges of the loads, because Enceladus doesn't check bounds on pointer accesses.
- Replace
triton.cdivwithenceladus.cdiv. Hints such asnum_stagescan stay; Enceladus ignores them. - Optional: Switch matmul loads to tensor descriptors, passing only the row strides.
What differs
The following table lists what differs:
| Area | Triton | Enceladus |
|---|---|---|
num_warps | Warps | SIMD groups of 32 threads, a power of two from 1 to 32 |
| Program | CTA | Metal threadgroup with 32 KB of threadgroup memory |
| Types | fp64, fp8, tf32 | No float64, FP8, or TF32. A float32 dot runs in full float32. |
tl.dot | Many dtypes, including integers | Same-dtype float16, bfloat16, or float32 operands. K must be a multiple of 8. |
| Descriptors | Needs a host allocator | No allocator. The innermost stride must be 1. |
| Atomics | Full sem and scope | Relaxed ordering only. No 16-bit float atomic_add; use a float32 buffer. |
| New launch options | None | dot_warps=(WM, WN) and dot_backend |
| Inline assembly | tl.inline_asm_elementwise | enceladus.metal_kernel(source, name) for hand-written MSL |
| Benchmarking | triton.testing.do_bench | enceladus.testing.do_bench |
Unsupported features
The following Triton features raise CompilationError:
- Block pointers:
tl.make_block_ptr,tl.advance,boundary_check, andpadding_option. while,break, andcontinue.tl.multiple_of,tl.max_contiguous,tl.join,tl.split,tl.dot_scaled,tl.sort,tl.flip,tl.gather,tl.histogram,tl.rand, and libdevice.**on runtime values. Usetl.exp2andtl.log2, or multiply.- Warp specialization, TMA, and clusters.