ai.onnx.TensorScatter

ai.onnx · standard ONNX operator · ONNX opset ≥ 24

Description

Functionally updates a KV cache tensor by scattering an update tensor into the past_cache along a sequence axis, producing present_cache with the same shape. Each batch sample's update is written at the offset given by write_indices (zero if omitted), either linearly or in wrap-around circular fashion.

See the ONNX TensorScatter spec for the reference semantics.

Inputs

Name Bind key Logical dtype WebGPU storage Rank Shape Description Presence
past_cache past T runtime-selected; narrow integers and bool use 32-bit slots Existing cache tensor with shape (batch_size, ..., max_sequence_length, ...). required
update update T runtime-selected; narrow integers and bool use 32-bit slots New values to scatter in, with the same shape as past_cache except the sequence dimension equals sequence_length. required
write_indices writeIndices I uint32 1 Logical int64 per-sample write offset into the cache sequence dimension; shape (batch_size,), stored as uint32 by WebGPU, and assumed all zeros if absent. optional

Outputs

Name Bind key Logical dtype Rank Shape Description Presence
present_cache present T same as past_cache same as past_cache Updated cache; same shape as past_cache. required

Attributes

Default values (overridable per request):

Attribute Default Description
axis -2 Sequence dimension of past_cache and update; cannot be 0 (the batch dimension). Default is -2.
mode "linear" Write mode: linear requires write_indices + sequence_length <= max_sequence_length; circular wraps the write index modulo max_sequence_length.

Type constraints

Variable Allowed dtypes
T float32, float16, int32, int16, int8, uint32, uint8, bool
I int64

Files

Use with @huggingface/kernels

The loader derives every required output's shape and logical dtype from the manifest contract and this call. It then allocates the result tensors automatically.

The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version.

Replace each *Data placeholder with a typed array containing the corresponding input data.

import { getKernel } from "@huggingface/kernels";

const kernel = await getKernel("webgpu-kernels/ai.onnx.TensorScatter", { version: 1 });
const { present } = await kernel({
  past: { data: pastData, shape: [1, 4, 2] },
  update: { data: updateData, shape: [1, 3, 2] },
});
Downloads last month
-
kernel
webgpu
wgsl
apache-2.0
WebGPU

Requires WebGPU support. See the compatibility table.