Skip to main content
Ctrl+K
AMD Logo
ROCm™ Software Ubuntu Build 7.2.4
Version List
  • GitHub
  • Community
  • Blogs
  • ROCm Developer Hub
  • ROCm Toolkits
    • ROCm Data Science
    • ROCm Finance
    • ROCm Life Science
    • ROCm Simulation
  • Systems and Infra Docs
  • Infinity Hub
  • Support

ROCm documentation

Composable Kernel 1.2.0 Documentation

Install

  • Install Composable Kernel
  • Build from source
  • Docker images

Conceptual

  • Structure
  • Mathematical basis
  • Intrawave and interwave scheduling
  • CK Tile conceptual documentation
    • CK Tile Conceptual Documentation
    • Introduction and Motivation - Why Tile Distribution Matters
    • Buffer Views - Raw Memory Access
    • Tensor Views - Multi-Dimensional Structure
    • Tile Distribution - The Core API
    • Coordinate Systems - The Mathematical Foundation
    • Terminology Reference - Key Concepts and Definitions
    • Tensor Adaptors - Chaining Transformations
    • Individual Transform Operations
    • Tensor Descriptors - Complete Tensor Specifications
    • Tile Window - Data Access Gateway
    • LoadStoreTraits - Memory Access Optimization Engine
    • Space-Filling Curves - Optimal Memory Traversal
    • Static Distributed Tensor
    • Convolution Implementation with CK Tile
    • Advanced Coordinate Movement
    • Load Data Share Index Swapping
    • Memory Swizzling with Morton Ordering
    • Tensor Coordinates
    • Sweep Tile
    • Encoding Internals
    • Thread Mapping - Connecting to Hardware
    • CK Tile Hardware Documentation
      • Intro to AMD CDNA Architecture
      • Understanding AMD GPU LDS and Bank Conflicts
      • A Block GEMM on MI300

Tutorial

  • Examples

Reference

  • Precision support
  • Scalar types
  • Custom types
  • Vector utilities
  • Wrapper
  • Glossary

About

  • Contributing to Composable Kernel
  • License

Index

A | B | C | D | E | F | G | H | I | K | L | M | N | O | P | R | S | T | U | V | W

A

  • Add+Multiply
  • alignment
  • arithmetic logic unit

B

  • bank conflict
  • batched GEMM
  • block Size
  • block tile

C

  • Col2Im
  • compute unit
  • coordinate transformation primitives

D

  • dense tensor
  • descriptor
  • device
  • dilation

E

  • elementwise
  • epilogue

F

  • fast changing dimension
  • fused add multiply

G

  • GEMM
  • GEMV
  • general matrix multiply
  • general matrix vector multiplication
  • GGEMM
  • global memory
  • grid
  • grouped GEMM

H

  • host
  • host-device transfer

I

  • Im2Col
  • inner dimension
  • input

K

  • kernel

L

  • launch parameters
  • LDS
  • LDS banks
  • load tile
  • local data share

M

  • matrix core
  • matrix fused multiply-add
  • memory coalescing
  • MFMA

N

  • naive GEMM

O

  • occupancy
  • operation
  • outer dimension

P

  • padding
  • permute
  • pinned memory
  • pipeline
  • policy
  • problem
  • problem shape

R

  • reference kernel
  • register

S

  • scalar general purpose register
  • SGPR
  • SIMD
  • SIMT
  • single-instruction, multi-data
  • single-instruction, multi-thread
  • sparse tensor
  • Split-K GEMM
  • store tile
  • stride

T

  • tile
  • tile distribution
  • tile partitioner
  • tile programming API
  • tile window
  • transpose

U

  • user customized tile pipeline
  • user customized tile pipeline optimization

V

  • vanilla GEMM
  • vector
  • vector general purpose register
  • VGEMM
  • VGPR

W

  • wave tile
  • wavefront
  • work group
  • work-item

  • Terms and Conditions
  • ROCm Licenses and Disclaimers
  • Privacy
  • Trademarks
  • Supply Chain Transparency
  • Fair and Open Competition
  • UK Tax Strategy
  • Cookie Policy
  • Cookie Settings
© 2026 Advanced Micro Devices, Inc