Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
2 changes: 1 addition & 1 deletion docs/arch/codegen.rst
Original file line number Diff line number Diff line change
Expand Up @@ -152,7 +152,7 @@ backend (x86, ARM, NVPTX, AMDGPU, etc.).
``Cast`` → LLVM type conversions, ``Call`` → intrinsic or extern function calls.
- **Statements** (``VisitStmt_``) emit LLVM IR side effects:
``BufferStore`` → store instructions, ``For`` → loop basic blocks with branches,
``IfThenElse`` → conditional branches, ``AllocBuffer`` → stack or heap allocation.
``IfThenElse`` → conditional branches, ``tirx.alloc_tensor`` calls → stack or heap allocation.

The key methods on ``CodeGenLLVM`` are:

Expand Down
2 changes: 1 addition & 1 deletion docs/arch/tvmscript.rst
Original file line number Diff line number Diff line change
Expand Up @@ -174,7 +174,7 @@ For example, a small function can be authored, printed and parsed again:
from tvm.script import tirx as T

@T.prim_func
def increment(A: T.Buffer((4,), "float32")):
def increment(A: T.Tensor((4,), "float32")):
for i in T.serial(4):
A[i] = A[i] + 1.0

Expand Down
6 changes: 3 additions & 3 deletions docs/deep_dive/relax/learning.rst
Original file line number Diff line number Diff line change
Expand Up @@ -133,13 +133,13 @@ for the end-to-end model execution. The code block below shows a TVMScript imple
class Module:
M, N, K = T.int64(), T.int64(), T.int64()
@Ts.prim_func(private=True)
def linear(X: T.Buffer((M, K), 'float32'), W: T.Buffer((K, N), 'float32'), B: T.Buffer((N,), 'float32'), Z: T.Buffer((M, N), 'float32')):
def linear(X: T.Tensor((M, K), 'float32'), W: T.Tensor((K, N), 'float32'), B: T.Tensor((N,), 'float32'), Z: T.Tensor((M, N), 'float32')):





Y = T.alloc_buffer((M, N), "float32")
Y = T.alloc_tensor((M, N), "float32")
for i, j, k in T.grid(M, N, K):
with Ts.sblock("Y"):
v_i, v_j, v_k = Ts.axis.remap("SSR", [i, j, k])
Expand All @@ -153,7 +153,7 @@ for the end-to-end model execution. The code block below shows a TVMScript imple

M, N = T.int64(), T.int64()
@Ts.prim_func(private=True)
def relu(X: T.Buffer((M, N), 'float32'), Y: T.Buffer((M, N), 'float32')):
def relu(X: T.Tensor((M, N), 'float32'), Y: T.Tensor((M, N), 'float32')):



Expand Down
10 changes: 5 additions & 5 deletions docs/deep_dive/relax/tutorials/relax_creation.py
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ def forward(
@I.ir_module
class RelaxModuleWithTIR:
@Ts.prim_func
def relu(X: T.Buffer((n, m), "float32"), Y: T.Buffer((n, m), "float32")):
def relu(X: T.Tensor((n, m), "float32"), Y: T.Tensor((n, m), "float32")):
for i, j in T.grid(n, m):
with Ts.sblock("relu"):
vi, vj = Ts.axis.remap("SS", [i, j])
Expand Down Expand Up @@ -170,10 +170,10 @@ def forward(self, x):

@Ts.prim_func
def tir_linear(
X: T.Buffer((M, K), "float32"),
W: T.Buffer((N, K), "float32"),
B: T.Buffer((N,), "float32"),
Z: T.Buffer((M, N), "float32"),
X: T.Tensor((M, K), "float32"),
W: T.Tensor((N, K), "float32"),
B: T.Tensor((N,), "float32"),
Z: T.Tensor((M, N), "float32"),
):
for i, j, k in T.grid(M, N, K):
with Ts.sblock("linear"):
Expand Down
6 changes: 3 additions & 3 deletions docs/deep_dive/tensor_ir/abstraction.rst
Original file line number Diff line number Diff line change
Expand Up @@ -34,9 +34,9 @@ the compute statements themselves.

@Ts.prim_func
def main(
A: T.Buffer((128,), "float32"),
B: T.Buffer((128,), "float32"),
C: T.Buffer((128,), "float32"),
A: T.Tensor((128,), "float32"),
B: T.Tensor((128,), "float32"),
C: T.Tensor((128,), "float32"),
) -> None:
for i in range(128):
with Ts.sblock("C"):
Expand Down
26 changes: 13 additions & 13 deletions docs/deep_dive/tensor_ir/learning.rst
Original file line number Diff line number Diff line change
Expand Up @@ -65,10 +65,10 @@ language called TVMScript, which is a domain-specific dialect embedded in python
@tvm.script.ir_module
class MyModule:
@Ts.prim_func
def mm_relu(A: T.Buffer((128, 128), "float32"),
B: T.Buffer((128, 128), "float32"),
C: T.Buffer((128, 128), "float32")):
Y = T.alloc_buffer((128, 128), dtype="float32")
def mm_relu(A: T.Tensor((128, 128), "float32"),
B: T.Tensor((128, 128), "float32"),
C: T.Tensor((128, 128), "float32")):
Y = T.alloc_tensor((128, 128), dtype="float32")
for i, j, k in T.grid(128, 128, 128):
with Ts.sblock("Y"):
vi = Ts.axis.spatial(128, i)
Expand All @@ -93,15 +93,15 @@ Function Parameters and Buffers
.. code:: python

# TensorIR
def mm_relu(A: T.Buffer((128, 128), "float32"),
B: T.Buffer((128, 128), "float32"),
C: T.Buffer((128, 128), "float32")):
def mm_relu(A: T.Tensor((128, 128), "float32"),
B: T.Tensor((128, 128), "float32"),
C: T.Tensor((128, 128), "float32")):
...
# NumPy
def lnumpy_mm_relu(A: np.ndarray, B: np.ndarray, C: np.ndarray):
...

Here ``A``, ``B``, and ``C`` takes a type named ``T.Buffer``, which with shape
Here ``A``, ``B``, and ``C`` takes a type named ``T.Tensor``, which with shape
argument ``(128, 128)`` and data type ``float32``. This additional information
helps possible MLC process to generate code that specializes in the shape and data
type.
Expand All @@ -111,7 +111,7 @@ type.
.. code:: python

# TensorIR
Y = T.alloc_buffer((128, 128), dtype="float32")
Y = T.alloc_tensor((128, 128), dtype="float32")
# NumPy
Y = np.empty((128, 128), dtype="float32")

Expand Down Expand Up @@ -240,10 +240,10 @@ So we can also write the programs as follows.
@tvm.script.ir_module
class MyModuleWithAxisRemapSugar:
@Ts.prim_func
def mm_relu(A: T.Buffer((128, 128), "float32"),
B: T.Buffer((128, 128), "float32"),
C: T.Buffer((128, 128), "float32")):
Y = T.alloc_buffer((128, 128), dtype="float32")
def mm_relu(A: T.Tensor((128, 128), "float32"),
B: T.Tensor((128, 128), "float32"),
C: T.Tensor((128, 128), "float32")):
Y = T.alloc_tensor((128, 128), dtype="float32")
for i, j, k in T.grid(128, 128, 128):
with Ts.sblock("Y"):
vi, vj, vk = Ts.axis.remap("SSR", [i, j, k])
Expand Down
28 changes: 14 additions & 14 deletions docs/deep_dive/tensor_ir/tutorials/tir_creation.py
Original file line number Diff line number Diff line change
Expand Up @@ -65,11 +65,11 @@
class MyModule:
@Ts.prim_func
def mm_relu(
A: T.Buffer((128, 128), "float32"),
B: T.Buffer((128, 128), "float32"),
C: T.Buffer((128, 128), "float32"),
A: T.Tensor((128, 128), "float32"),
B: T.Tensor((128, 128), "float32"),
C: T.Tensor((128, 128), "float32"),
):
Y = T.alloc_buffer((128, 128), dtype="float32")
Y = T.alloc_tensor((128, 128), dtype="float32")
for i in range(128):
for j in range(128):
for k in range(128):
Expand Down Expand Up @@ -108,11 +108,11 @@ def mm_relu(
class ConciseModule:
@Ts.prim_func
def mm_relu(
A: T.Buffer((128, 128), "float32"),
B: T.Buffer((128, 128), "float32"),
C: T.Buffer((128, 128), "float32"),
A: T.Tensor((128, 128), "float32"),
B: T.Tensor((128, 128), "float32"),
C: T.Tensor((128, 128), "float32"),
):
Y = T.alloc_buffer((128, 128), dtype="float32")
Y = T.alloc_tensor((128, 128), dtype="float32")
for i, j, k in T.grid(128, 128, 128):
with Ts.sblock("Y"):
vi, vj, vk = Ts.axis.remap("SSR", [i, j, k])
Expand Down Expand Up @@ -147,11 +147,11 @@ def mm_relu(
class ConciseModuleFromPython:
@Ts.prim_func
def mm_relu(
A: T.Buffer((M, K), dtype),
B: T.Buffer((K, N), dtype),
C: T.Buffer((M, N), dtype),
A: T.Tensor((M, K), dtype),
B: T.Tensor((K, N), dtype),
C: T.Tensor((M, N), dtype),
):
Y = T.alloc_buffer((M, N), dtype)
Y = T.alloc_tensor((M, N), dtype)
for i, j, k in T.grid(M, N, K):
with Ts.sblock("Y"):
vi, vj, vk = Ts.axis.remap("SSR", [i, j, k])
Expand Down Expand Up @@ -185,10 +185,10 @@ def mm_relu(
@I.ir_module
class DynamicShapeModule:
@Ts.prim_func
def mm_relu(A: T.Buffer([M, K], dtype), B: T.Buffer([K, N], dtype), C: T.Buffer([M, N], dtype)):
def mm_relu(A: T.Tensor([M, K], dtype), B: T.Tensor([K, N], dtype), C: T.Tensor([M, N], dtype)):
# Bind the input buffers with the dynamic shapes

Y = T.alloc_buffer((M, N), dtype)
Y = T.alloc_tensor((M, N), dtype)
for i, j, k in T.grid(M, N, K):
with Ts.sblock("Y"):
vi, vj, vk = Ts.axis.remap("SSR", [i, j, k])
Expand Down
6 changes: 3 additions & 3 deletions docs/deep_dive/tensor_ir/tutorials/tir_transformation.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,9 +46,9 @@
class MyModule:
@Ts.prim_func
def main(
A: T.Buffer((128, 128), "float32"),
B: T.Buffer((128, 128), "float32"),
C: T.Buffer((128, 128), "float32"),
A: T.Tensor((128, 128), "float32"),
B: T.Tensor((128, 128), "float32"),
C: T.Tensor((128, 128), "float32"),
):
T.func_attr({"tirx.noalias": True})
with Ts.sblock("root"):
Expand Down
26 changes: 13 additions & 13 deletions docs/how_to/tutorials/mix_python_and_tvm_with_pymodule.py
Original file line number Diff line number Diff line change
Expand Up @@ -87,9 +87,9 @@
class MyFirstModule(BasePyModule):
@Ts.prim_func
def add_tir(
A: T.Buffer((4,), "float32"),
B: T.Buffer((4,), "float32"),
C: T.Buffer((4,), "float32"),
A: T.Tensor((4,), "float32"),
B: T.Tensor((4,), "float32"),
C: T.Tensor((4,), "float32"),
):
for i in range(4):
C[i] = A[i] + B[i]
Expand Down Expand Up @@ -133,9 +133,9 @@ def forward(self, x, y):
class DebugModule(BasePyModule):
@Ts.prim_func
def matmul_tir(
A: T.Buffer((n, 4), "float32"),
B: T.Buffer((4, 3), "float32"),
C: T.Buffer((n, 3), "float32"),
A: T.Tensor((n, 4), "float32"),
B: T.Tensor((4, 3), "float32"),
C: T.Tensor((n, 3), "float32"),
):
for i, j, k in T.grid(n, 3, 4):
with Ts.sblock("matmul"):
Expand Down Expand Up @@ -210,9 +210,9 @@ def my_bias_add(x, bias, out):
class PipelineModule(BasePyModule):
@Ts.prim_func
def matmul_tir(
A: T.Buffer((2, 4), "float32"),
B: T.Buffer((4, 3), "float32"),
C: T.Buffer((2, 3), "float32"),
A: T.Tensor((2, 4), "float32"),
B: T.Tensor((4, 3), "float32"),
C: T.Tensor((2, 3), "float32"),
):
for i, j, k in T.grid(2, 3, 4):
with Ts.sblock("matmul"):
Expand Down Expand Up @@ -274,9 +274,9 @@ def forward(self, x, weights, bias):
class DenseLayer:
@Ts.prim_func
def bias_add_tir(
x: T.Buffer((2, 4), "float32"),
b: T.Buffer((4,), "float32"),
out: T.Buffer((2, 4), "float32"),
x: T.Tensor((2, 4), "float32"),
b: T.Tensor((4,), "float32"),
out: T.Tensor((2, 4), "float32"),
):
for i, j in T.grid(2, 4):
out[i, j] = x[i, j] + b[j]
Expand Down Expand Up @@ -401,7 +401,7 @@ def main(
@R.py_module
class DynamicModule(BasePyModule):
@Ts.prim_func
def scale_tir(x: T.Buffer((n,), "float32"), out: T.Buffer((n,), "float32")):
def scale_tir(x: T.Tensor((n,), "float32"), out: T.Tensor((n,), "float32")):
for i in T.serial(n):
out[i] = x[i] * T.float32(2.0)

Expand Down
2 changes: 1 addition & 1 deletion docs/tirx/api/script.rst
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ TIRx kernels use ``tvm.script.tirx`` for the parser and core IR builders::

from tvm.script import tirx as Tx

Tx.alloc_buffer(...)
Tx.alloc_tensor(...)

Tile primitives and backend-specific namespaces are documented separately in
:doc:`tile`, :doc:`cuda`, and :doc:`ptx`. For the relationship between these
Expand Down
2 changes: 1 addition & 1 deletion docs/tirx/api/tirx.rst
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ here so the same objects are not expanded twice.

For C++ construction, include ``tvm/tirx/expr.h`` for ``BufferVar``, buffer
loads, and buffer-region constructors. Include ``tvm/tirx/type.h`` for
``BufferType``, ``BufferRegionType``, and ``TensorMapType``. Buffer regions
``TensorType``, ``BufferRegionType``, and ``TensorMapType``. Buffer regions
use the shared ``TensorRegion`` expression from ``tvm/ir/expr.h``.

.. automodule:: tvm.tirx
Expand Down
2 changes: 1 addition & 1 deletion docs/tirx/arch/lowering_pipeline.rst
Original file line number Diff line number Diff line change
Expand Up @@ -157,7 +157,7 @@ Take a one-line scale kernel:
.. code-block:: python

@Tx.prim_func
def scale(A: Tx.Buffer((256,), "float32"), B: Tx.Buffer((256,), "float32")):
def scale(A: Tx.Tensor((256,), "float32"), B: Tx.Tensor((256,), "float32")):

Tx.device_entry()
bx = Tx.cta_id([1])
Expand Down
2 changes: 1 addition & 1 deletion docs/tirx/native_basics.rst
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ The authoring model
- ``@Tx.prim_func`` (or ``@Tx.jit`` for compile-time-specialized) kernels, written
with ``from tvm.script import tirx as Tx``;
- ``Tx.device_entry()`` plus *scope-id* intrinsics for thread binding;
- ``Tx.Buffer`` parameter annotations and ``Tx.alloc_*`` scratch buffers;
- ``Tx.Tensor`` parameter annotations and ``Tx.alloc_*`` scratch buffers;
- ordinary loops, branches, and scalar math;
- ``tvm.compile(mod, target=..., tir_pipeline="tirx")`` to build, then call the
result directly.
Expand Down
Loading
Loading