ai.onnx.TensorScatter
ai.onnx · standard ONNX operator · ONNX opset ≥ 24
Description
Functionally updates a KV cache tensor by scattering an update tensor into the past_cache along a sequence axis, producing present_cache with the same shape. Each batch sample's update is written at the offset given by write_indices (zero if omitted), either linearly or in wrap-around circular fashion.
See the ONNX TensorScatter spec for the reference semantics.
Inputs
| Name | Bind key | Logical dtype | WebGPU storage | Rank | Shape | Description | Presence |
|---|---|---|---|---|---|---|---|
past_cache |
past |
T |
runtime-selected; narrow integers and bool use 32-bit slots | — | — | Existing cache tensor with shape (batch_size, ..., max_sequence_length, ...). |
required |
update |
update |
T |
runtime-selected; narrow integers and bool use 32-bit slots | — | — | New values to scatter in, with the same shape as past_cache except the sequence dimension equals sequence_length. |
required |
write_indices |
writeIndices |
I |
uint32 |
1 |
— | Logical int64 per-sample write offset into the cache sequence dimension; shape (batch_size,), stored as uint32 by WebGPU, and assumed all zeros if absent. |
optional |
Outputs
| Name | Bind key | Logical dtype | Rank | Shape | Description | Presence |
|---|---|---|---|---|---|---|
present_cache |
present |
T |
same as past_cache |
same as past_cache |
Updated cache; same shape as past_cache. |
required |
Attributes
Default values (overridable per request):
| Attribute | Default | Description |
|---|---|---|
axis |
-2 |
Sequence dimension of past_cache and update; cannot be 0 (the batch dimension). Default is -2. |
mode |
"linear" |
Write mode: linear requires write_indices + sequence_length <= max_sequence_length; circular wraps the write index modulo max_sequence_length. |
Type constraints
| Variable | Allowed dtypes |
|---|---|
T |
float32, float16, int32, int16, int8, uint32, uint8, bool |
I |
int64 |
Files
metadata.json— kernel metadata (id, digests, provenance)manifest.json— the op contract (source of truth)test.json— correctness casesbench.json— benchmark + tuning casesscatter-flat-copy.wgsl.jinjatensor-scatter.wgsl.jinja
Use with @huggingface/kernels
The loader derives every required output's shape and logical dtype from the manifest contract and this call. It then allocates the result tensors automatically.
The version: 1 option selects the published kernel contract; it is independent of any operator opset, contrib since_version, or model version.
Replace each *Data placeholder with a typed array containing the corresponding input data.
import { getKernel } from "@huggingface/kernels";
const kernel = await getKernel("webgpu-kernels/ai.onnx.TensorScatter", { version: 1 });
const { present } = await kernel({
past: { data: pastData, shape: [1, 4, 2] },
update: { data: updateData, shape: [1, 3, 2] },
});
- Downloads last month
- -
Requires WebGPU support. See the compatibility table.